Sunday, 16 August 2026
The Verified Journalism Press

Journalism with its sources attached.

Sections
WORLD
AUSTRALIA
INDIA
BUSINESS
TECHNOLOGY
SCIENCE
SOCIETY
RIGHTS
CORRUPTION
CULTURE
OPINION
FAMOUS
The Press
Latest
Brussels has child safety cases open against Snapchat, Meta and TikTok, but not YouTube or the app storesMost Australian under-16s are still using social media, the regulator's own evaluation findsAI-designed viruses clear peer review, then an independent check finds them close relatives of the natural originalMIT's AI supercomputer has fallen 36 places in the world rankings without getting any slowerArizona physicists shift the quantum noise inside a light pulse, and watch it move in real timeApple has handed Siri to Google, and Amazon's Alexa+ has reached AustraliaBrussels has child safety cases open against Snapchat, Meta and TikTok, but not YouTube or the app storesMost Australian under-16s are still using social media, the regulator's own evaluation findsAI-designed viruses clear peer review, then an independent check finds them close relatives of the natural originalMIT's AI supercomputer has fallen 36 places in the world rankings without getting any slowerArizona physicists shift the quantum noise inside a light pulse, and watch it move in real timeApple has handed Siri to Google, and Amazon's Alexa+ has reached Australia
Markets
ASX 200
S&P 500
Nasdaq
FTSE 100
Nikkei
Gold
Brent
AUD / USD
AUD / EUR
AUD / GBP
AUD / JPY
Bitcoin
Ethereum
Yahoo · ECB · CoinGecko

Front page / Artificial Intelligence

Copyright and courts

A judge made OpenAI hand over 20 million ChatGPT conversations, then ordered 88 million more

Judge Sidney Stein affirmed an order compelling OpenAI to produce a full 20 million log sample to the copyright plaintiffs on 5 January 2026. What OpenAI lost was not the sample size, which it proposed itself, but its attempt to limit production to conversations touching the plaintiffs' works.

Daniel Patrick Moynihan United States Courthouse (55314001073)
Daniel Patrick Moynihan United States Courthouse (55314001073). Photograph: 83136374@N05, CC BY 4.0

The widely repeated version of this story is that a federal judge forced OpenAI to surrender 20 million private ChatGPT conversations over the company's objections. The number is right and the order is real, but the 20 million figure was OpenAI's own proposal, not the court's imposition. What the company lost was the argument about what would be inside it.

In re: OpenAI, Inc. Copyright Infringement Litigation is a multidistrict proceeding in the Southern District of New York consolidating sixteen copyright suits brought by news organisations and authors, among them The New York Times and the Chicago Tribune. Plaintiffs originally sought 120 million chat logs. OpenAI countered with 20 million de identified logs, and the plaintiffs accepted that sample size while reserving the right to ask for more. OpenAI then sought to narrow what the sample contained, proposing to hand over only conversations that implicated the plaintiffs' specific works and arguing that logs which did not contain those works were irrelevant.

Magistrate Judge Ona T. Wang rejected that approach in November 2025. District Judge Sidney Stein affirmed. Writing for the National Law Review on 6 January 2026, Andrew R. Lee of Jones Walker dates the affirmance to 5 January 2026 and reports that the court found even unrelated logs discoverable because they bear on OpenAI's fair use defence, and specifically on the effect of the use upon the market for the copyrighted works. The ABA Journal, publishing on 8 January 2026, reported the same ruling and recorded Judge Stein's finding that "Judge Wang's rulings were neither clearly erroneous nor contrary to law". The ABA Journal gives no calendar date, saying only that the judge ruled on the Monday, which that week fell on 5 January. The two accounts agree.

The privacy reasoning is the part with consequences beyond this case. Judge Stein identified three protections as sufficient: the reduction of the sample from billions of conversations to 20 million, OpenAI's own de identification stripping personally identifiable and other private information, and the protective order already governing discovery material. On OpenAI's argument that production would invade its users' privacy, the court distinguished ChatGPT users from the subjects of a wiretap on the basis that users had voluntarily submitted their communications to the company.

It did not stop at 20 million. Norton Rose Fulbright, in a March 2026 update on AI copyright litigation, records the 5 January order requiring OpenAI to produce all 20 million output logs, quoting the court's observation that "OpenAI identifies no caselaw requiring a court to order the least burdensome discovery possible or to explain specifically why it rejects a party's discovery proposal". The same firm records that on 9 March 2026 the court granted the plaintiffs' motion to compel and ordered OpenAI to produce "the reservoirs of 78 million and 10 million logs", on top of the 20 million already ordered. That is 108 million conversations in total, five times the number in the headlines, and this masthead has found no general news coverage of the March order comparable to the coverage of the January one.

Nothing has yet been decided on the merits. The AI copyright tracker published by Axis Intelligence records the 20 million logs as being in evidence and fair use briefing as closed on 2 April 2026, with no ruling. Writing in the National Law Review on 13 August 2026, Adam Eisgrau reported that the consolidated case is expected to be argued early in 2027 before a federal judge in New York, and that the plaintiffs, described as dozens of authors including George R. R. Martin along with major news organisations, are seeking orders that would require existing models to be withdrawn and retrained on licensed material only.

What remains unknown is what de identification actually removed. None of the published accounts describe the technique, its error rate, or whether anyone has tested whether a conversation stripped of names can still be traced to the person who typed it. The court accepted de identification as adequate; no source reviewed here shows it was independently audited.

Sources

Every factual claim above rests on the 5 published sources below. They are listed so you can check the reporting rather than take it on trust.

  1. ABA JournalChatGPT creator must turn over 20M chat logs in copyright litigation, federal judge says
  2. National Law Review (Andrew R. Lee, Jones Walker LLP)OpenAI Loses Privacy Gambit: 20 Million ChatGPT Logs Likely Headed to Copyright Plaintiffs
  3. Norton Rose FulbrightAI in litigation series: an update on AI copyright cases in 2026
  4. National Law Review (Adam Eisgrau)The Copyright Battle That Could Push AI Out of Reach
  5. Axis IntelligenceAI Copyright Lawsuits Status Tracker

The Verified Briefing

One email each morning. Every story in it carries its sources, so you can check the reporting before you repeat it.

No tracking pixels. One click to leave.