NYT v. OpenAI: how to read an evidence accusation without mistaking it for a ruling
Publishers accuse OpenAI of concealing searches and losing logs. A guide to separating allegations, prior rulings and sanctions that have only been requested.
On July 9, 2026, The New York Times and other publishers asked a federal court in New York to sanction OpenAI for its conduct during discovery in the copyright litigation. The important news is not that a judge concluded OpenAI had concealed or destroyed evidence: no such ruling existed on that date. What existed was a 52-page motion filed by the plaintiffs, containing serious allegations and five requested remedies, alongside earlier orders that let a reader verify which parts of the story the court had already established.
That distinction may sound formal, but it changes everything a reader can responsibly say. A motion presents the account of the party asking for relief; its exhibits provide the material offered to support that account; an order states what the court decided; and a judgment resolves a dispute. Blending those levels turns an accusation into a proven fact. The durable skill in this case is learning to read any legal story through four questions: who is speaking, which document contains the statement, what has the judge already decided, and what remedy is merely being requested?
The allegation in the original filing
The publishers’ memorandum says OpenAI represented during the case that it lacked efficient tools to search for particular works in its training datasets and lacked infrastructure to look for specific content across large collections of conversations. According to the plaintiffs, a second deposition of Vincent Monaco, an OpenAI corporate representative, held on April 8, 2026, revealed internal capabilities and searches that contradicted those positions.
The filing divides its account among three types of material. The first consists of training datasets: the publishers say OpenAI had already searched them for publisher content. The second consists of ChatGPT conversation reservoirs: they allege that by June 2023 a set of roughly 78 million de-identified, searchable conversations existed. The third is Project Giraffe, a name covering tools and evaluations used to study content reproduction. Specific passages in the public memorandum are redacted, so the document can verify that these allegations were made but cannot reveal the entire technical operation.
That limitation blocks two shortcuts. A Bloom filter is a probabilistic data structure that can help test whether an item may belong to a set; by itself it does not decide whether copyright infringement occurred. The existence of an internal evaluation for reproduction likewise does not automatically prove that every detected output was unlawful. It may be relevant evidence about what the company could search and when, but the legal consequence depends on the content, the use, the defenses and the court’s evaluation.
What the court had already determined
Some procedural facts do not rest solely on the publishers’ account. In a March 9, 2026 order, Magistrate Judge Ona T. Wang directed production of reservoirs containing 78 million and 10 million logs under a de-identification protocol. The same order found that Monaco had not been sufficiently or properly prepared for his first corporate deposition and that the pattern of objections and resulting answers impeded, delayed and frustrated a fair examination. It also directed the parties to discuss discovery remedies and the production of non-privileged Project Giraffe documents.
That order carries more weight than a journalistic paraphrase, but it still does not say everything sometimes attributed to it. It does not find that Project Giraffe proved infringement. It does not decide the new sanctions request. Indeed, it expressly leaves spoliation claims for separate briefing. A careful reading preserves the court’s verbs: it ordered production, found deficiencies in a deposition, and reserved the sanctions issue for later.
Another judicial document establishes the preservation baseline. On May 13, 2025, the court ordered OpenAI to preserve and segregate, from that point forward and until further order, all output logs that would otherwise be deleted. The order also acknowledged OpenAI’s stated conflict involving deletion requests and privacy laws. A precise account must therefore separate deletion before the order, the broader preservation duties the plaintiffs say had already arisen, and compliance after the explicit command. The motion incorporates all those periods into its theory; the court still had to assess each one.
The 20-million sample and representativeness
The publishers say they spent months negotiating a sample of 20 million conversations because OpenAI described broader searches as technically expensive. They allege that the company kept deleting or compressing billions of conversations, replaced approximately 10 percent of the records first selected, and produced a version so heavily redacted that the court called it unusable. The memorandum says a less-redacted version arrived one week before fact discovery closed.
The transferable concept here is not the spectacular number but selection bias. A sample represents the universe from which it is drawn only if the inclusion mechanism does not systematically remove relevant cases. When records disappear because of retention policies, user requests or technical substitutions, the reader should ask which population remained available and whether the missing cases relate to the conduct being studied. Twenty million can sound enormous and still form a defective sample. Size reduces random error; it does not repair biased selection.
That point does not establish the publishers’ theory in advance. They must connect missing information to material that should have been preserved and demonstrate prejudice. OpenAI can contest the description of its systems, the prospects of restoring data, the proportionality of the requests, the actual effect on the sample and the intent attributed to it. As of July 13, the public memorandum was an opening submission, not a complete adversarial record. The fact that a formal opposition was not yet available did not amount to an admission.
What the rules permit and what the publishers want
Federal Rule of Civil Procedure 37 supplies different remedies for different failures. Rule 37(b) addresses disobedience of a discovery order and can allow a court to treat designated facts as established or prevent a party from supporting specified defenses. Rule 37(e) addresses electronically stored information that should have been preserved but was lost because reasonable steps were not taken. If the loss causes prejudice, the court may impose measures sufficient to cure it; the harsher presumptions tied to unfavorable information require a finding that a party intended to deprive the opponent of its use.
The publishers requested five forms of relief. First, OpenAI would be barred from relying on the 20-million-log sample for any purpose. Second, the court would find that the output logs contain — or would have shown if properly produced — substantial and systematic grounding on and regurgitation of publisher material. Third, OpenAI would be prohibited from arguing the contrary. Fourth, the jury would be instructed about those findings and their binding effect. Fifth, OpenAI would pay the publishers’ attorney and expert fees and the costs of pursuing the allegedly withheld material.
Those remedies could reshape a trial, which is precisely why they must not be described as granted. The court must determine which rule applies to each act, whether an order was violated, what information was lost, whether it can be restored, what prejudice resulted and, for particular sanctions, the relevant intent. It could also impose a narrower measure, reject parts of the request or distinguish among separate problems. “The plaintiffs seek an inference” and “the judge found the fact” are not interchangeable statements.
A method for resisting the headline
When the next technology lawsuit generates news, open the document and label every claim. “The plaintiff alleges” belongs in the position column. “The exhibit shows” requires checking that the exhibit is public and supports the exact proposition. “The court orders” belongs to a ruling and should carry its date and scope. “The law permits” describes a menu, not the outcome. Finally, look for the opposing brief: when it does not yet exist, make that absence explicit.
Applied here, that method produces a calmer and more useful account. On July 9 the publishers made a documented accusation; an earlier order confirms the existence of 78-million and 10-million reservoirs, problems with Monaco’s first deposition and a genuine Project Giraffe dispute; another order imposed prospective preservation; and Rule 37 authorizes sanctions when conditions the judge must determine are met. What did not exist on July 13 was a ruling on this motion. A reader who preserves that boundary can explain the dispute without accidentally becoming counsel for either side.
Sources for this piece
This piece draws on 4 primary source(s), gathered during reporting.
- Docket CourtListener/RECAP — NYT v. Microsoft/OpenAI, 1:23-cv-11195 (S.D.N.Y.) + MDL 25-md-3143
- Memorando de la moción de sanciones (MDL Dkt. 1618 / Times ECF 1427-1, 09/07/2026) — texto íntegro público
- TechCrunch (09/07/2026) — respuesta oficial de OpenAI a la moción
- Marco jurídico para el explainer (spoliation, Regla 37, adverse inference) + verdad de época
This article was produced with artificial intelligence under human editorial oversight.