An order no customer saw coming

On 13 May 2025, in the consolidated copyright cases against OpenAI in the Southern District of New York, Magistrate Judge Ona Wang directed OpenAI to "preserve and segregate all output log data that would otherwise be deleted on a going forward basis". The order applied whether the data "might be deleted at a user's request or because of 'numerous privacy laws and regulations'". OpenAI objected that it would have to disregard commitments to "hundreds of millions of people, businesses, educational institutions, and governments." Reconsideration was denied, and the district judge overruled OpenAI's objection.

The people whose conversations were affected were not parties. Neither were the businesses using the API. They had agreed retention terms with a vendor, and a court that had never heard of them set those terms aside.

What actually happened, precisely

Precision matters here, because the pitch version of this story is usually wrong.

ChatGPT Enterprise was excluded. OpenAI says so, and filings from both sides record that the court dropped it at a May 2025 hearing. OpenAI also says zero-data-retention API traffic was never stored, so the order could not reach it. Ordinary API traffic was different. OpenAI later told the court it "preserved millions of API logs" under the order.

The ongoing obligation ended on 26 September 2025. Data already segregated stays preserved, except for requests from the EEA, Switzerland and the UK, and OpenAI keeps preserving logs for accounts on a list of plaintiff-named domains. Separately, the court ordered production of 20 million de-identified consumer conversations under an attorneys'-eyes-only protective order. The district judge affirmed that order in January 2026. A sanctions motion about deleted logs is fully briefed and undecided. Nobody has been found to have done anything wrong.

The part that should worry a compliance officer

The sentence worth reading twice comes from the January 2026 affirmance. The court weighed the privacy of conversations "which users voluntarily disclosed to OpenAI and which OpenAI retains in the normal course of its business" and found it weaker than the privacy interest in wiretap recordings. Earlier, the magistrate judge wrote that privacy "cannot predominate where there is clear relevance and minimal burden."

None of this is unusual law. Federal Rule 34 reaches anything in a party's "possession, custody, or control." Rule 45 lets a subpoena reach non-parties the same way. Privacy is one factor in proportionality, not a veto. OpenAI is not alone either: courts have ordered Anthropic to produce millions of prompt-output pairs, and Midjourney has agreed to produce subscriber prompts. Those were production orders, not orders to override deletion, but they show the same thing. A vendor's log store is discoverable in litigation that has nothing to do with you.

For a bank, the retention schedule is a regulatory artefact. For a law firm, a promise to destroy a client's files is an ethical obligation. For a hospital, it may be both. Each can be overridden by a docket entry in a case the firm does not know exists, against a vendor it chose for other reasons.

Where the logs sit decides who answers

Running inference inside your own environment does not put data beyond discovery. Rule 34 reaches your own records too, and it should. What changes is the route. When the logs are yours, a preservation demand has to come to you and your counsel. You see it, you negotiate its scope, and you apply your own legal hold with your own retention policy intact everywhere else. When a vendor holds them, you may learn about a hold from a blog post.

That is a structural point, not a criticism of any one company. Every hosted provider is exposed to the same rules, however good its intentions. Vendors are responding with enterprise exclusions, zero-retention tiers and regional carve-outs, and these help. But they are negotiated and then litigated, one case at a time, by someone else.

The question to put to any AI vendor in a security review is not only "do you delete my data?" It is "who can make you stop, and would I find out?" For work that carries a duty of confidentiality, the best answer is one you give yourself.

Frequently asked questions

Is the OpenAI preservation order still in force?

No. The ongoing obligation ended on 26 September 2025. Data already preserved before then remains held, except requests from the EEA, Switzerland and the UK, and logs for certain plaintiff-named domains continue to be preserved.

Were enterprise customers' logs preserved?

ChatGPT Enterprise was excluded, and OpenAI says zero-data-retention API traffic was never stored. Standard API traffic was preserved during the order.

Does running AI on-premise protect data from discovery?

No. Your own records remain discoverable. What changes is that preservation and production demands come to you and your counsel, rather than to a vendor you may not hear from.

Sources & further reading

Talk with us about your workflow →