top of page

NYT vs. OpenAI

Jul 29
5 min read

Two and a half years after The New York Times became the first major news organization to sue OpenAI, the case has evolved from a straightforward copyright dispute into something closer to a forensic investigation of how ChatGPT actually works behind the scenes, and a fight over whether OpenAI has been telling the truth to a federal court.

How It Started

The Times filed suit against OpenAI and Microsoft in December 2023 in the Southern District of New York, alleging the companies used millions of Times articles without a license to train GPT models, and that ChatGPT could reproduce large chunks of Times journalism nearly verbatim, effectively substituting for the product the Times sells. The suit named Microsoft as a contributory infringer for supplying the computing infrastructure and business backing that made OpenAI's training possible.

The case was eventually consolidated with more than a dozen similar suits, from the Daily News, Chicago Tribune, Center for Investigative Reporting, Ziff Davis, Raw Story, AlterNet, and a group of authors, into a multidistrict proceeding, In re: OpenAI, Inc. Copyright Infringement Litigation, before Judge Sidney H. Stein, with discovery overseen by Magistrate Judge Ona T. Wang.


In March 2025, Judge Stein largely denied OpenAI and Microsoft's motions to dismiss. The Times' core direct and contributory copyright infringement claims were allowed to proceed; some peripheral claims, like DMCA notice-removal claims, common-law unfair competition, and abridgment claims from the Center for Investigative Reporting, were narrowed or dismissed. Crucially, Stein declined to rule on fair use at this stage, finding the question depended on facts that discovery hadn't yet developed, such as how the training data was used, whether outputs actually compete with Times journalism, and what the economic impact really is.

That set the stage for what has become the real battleground: getting inside ChatGPT's logs.



In May 2025, Magistrate Judge Wang issued a sweeping preservation order requiring OpenAI to retain all ChatGPT conversation logs going forward, covering more than 400 million users worldwide, after the Times accused OpenAI of deleting conversations that might show infringement. Judge Stein affirmed it in June 2025, rejecting OpenAI's argument that the order trampled user privacy. The mass-preservation requirement for new content was eventually lifted in September 2025, but by then the plaintiffs had already begun combing through the retained data.

The publishers then sought a sample of 120 million ChatGPT logs; the two sides negotiated that down to 20 million. When OpenAI later tried to limit what it would actually hand over, Judge Wang sided with the news organizations in November 2025, ordering the entire de-identified 20-million-log sample produced, including logs with no obvious connection to Times content, because they could still bear on OpenAI's fair-use defense. Judge Stein affirmed that ruling on January 5, 2026, calling it neither clearly erroneous nor contrary to law. In March 2026, Wang went further still, ordering OpenAI to produce 88 million pre-filter output logs and even portions of an OpenAI executive's personal journal.


The Deposition

The case took its sharpest turn this spring. After the court found OpenAI's corporate witness unprepared at an earlier deposition, it ordered a re-designated witness, OpenAI data privacy engineer Vincent Monaco, to testify again in April 2026. According to the plaintiffs, Monaco's testimony blew a hole in two years of OpenAI's own representations to the court: that OpenAI had, in fact, already searched its ChatGPT logs for copyrighted journalism before the Times even filed suit, had built an internal database of roughly 78 million de-identified conversations for that purpose, and had deployed a detection system (reportedly called Project Giraffe, using a Bloom filter) to flag regurgitated content shortly after the lawsuit began.

That directly contradicted OpenAI's long-standing argument in court that searching its logs for copyrighted material was technically infeasible" and unduly invasive of user privacy.


Sanctions Motion

On July 9, 2026, the Times, the Daily News, and other publisher plaintiffs filed a motion asking Judge Stein to sanction OpenAI, accusing the company of lying to the court for over two years about its ability to search ChatGPT outputs, deleting billions of conversations after the lawsuit was filed in violation of the preservation order, and turning over a 20-million-log sample so heavily redacted that the plaintiffs call it unusable.

The requested sanctions are aggressive: barring OpenAI from relying on the 20-million-log sample for any purpose, and asking the court to simply find that ChatGPT's output logs would have shown infringement of the plaintiffs' works had they been properly produced.

OpenAI, for its part, maintains its data-handling decisions were about protecting user privacy, not concealing evidence, and has characterized the Times' escalating discovery demands as an invasion of the privacy of people uninvolved in the case. The company has not yet filed a formal response to the sanctions motion.



Just two weeks before the sanctions motion, on June 25, 2026, the Times filed a third amended complaint, sharpening its allegation that Microsoft actively encouraged and enabled OpenAI's use of Times content, while voluntarily dropping a secondary-infringement claim against OpenAI over user-generated outputs. A Times spokesperson framed it as narrowing the case to its strongest arguments rather than a retreat; OpenAI has pointed to the dropped claim as evidence the case is weakening.

Where Things Stand

As of late July 2026:

  • Fair use remains unresolved. Judge Stein has repeatedly declined to decide it, and summary judgment briefing was scheduled to close in early April 2026, meaning a ruling could come at any time, now complicated by the sanctions fight.

  • No trial date has been set.

  • The sanctions motion is pending, with outcomes ranging from evidentiary penalties to a potentially case-defining finding that OpenAI's logs demonstrate infringement.

  • OpenAI has publicly pointed to favorable transformative-use rulings in other AI copyright cases (including the Anthropic and Meta author lawsuits) as precedent it hopes will carry over here, even as those very cases have produced mixed, piracy-sensitive outcomes rather than a clean industry-wide win.

What began as a dispute over article licensing has become a closely watched test of how courts will handle discovery into AI systems' internal workings and now, whether an AI company can be found to have misled a federal court about what it knew all along.


Read more here:

 
 
 

Recent Posts

See All
Sony and Warner vs. Anthropic

The music industry’s next major battle over artificial intelligence is underway. Sony Music Publishing, Warner Chappell Music, and 33 affiliated music publishers have sued Anthropic, alleging that the

 
 
 
Dolly Parton's Death and the Industry She Built

Dolly Parton died on Tuesday, August 25, 2026, at the Vanderbilt-Ingram Cancer Center in Nashville, surrounded by loved ones. She was 80 years old, and her publicist, Marcel Pariseau, said the cause w

 
 
 
Spotify's War on "AI Slop"

If you've seen headlines this month claiming Spotify deleted 75 million AI songs, the real story is both bigger and more nuanced than that number suggests. Spotify isn't banning AI music, it's runnin

 
 
 

Comments


bottom of page