OpenAI Is Hiding the Truth About Pirated Book Datasets—Why Were They Deleted?

OpenAI Is Hiding the Truth About Pirated Book Datasets—Why Were They Deleted?

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
3. 12. 2025
4 minutes reading · 13 views
OpenAI Is Hiding the Truth About Pirated Book Datasets—Why Were They Deleted?

OpenAI has found itself in serious trouble. At the center of attention is a class-action lawsuit filed by book authors who claim that their works were illegally used to train artificial intelligence. The key issue? OpenAI deleted two datasets full of pirated books, called Books1 and Books2, before launching ChatGPT in 2022. These datasets were created by former OpenAI employees in 2021 by crawling the open web and stealing most of the data from a so-called shadow library known as Library Genesis, or LibGen for short. OpenAI claims that it stopped using the datasets that same year and decided internally to delete them. But the authors believe there is more to the story, and it now appears that the court may agree with them.

Last week, U.S. Magistrate Judge Ona Wang ruled that OpenAI must share all internal communications with the company's lawyers regarding the deletion of these datasets. It must also provide all internal references to LibGen that OpenAI had withheld under the guise of attorney-client privilege. Wang ordered OpenAI to turn over these materials by December 8 and make its in-house lawyers available for questioning by December 19. Why such a ruling? According to the judge, OpenAI acted improperly—first claiming that the datasets being "unused" was the reason for deleting them, then withdrawing that claim and suddenly declaring all reasons privileged. Wang described this as "vacillation" and said that OpenAI cannot shift its positions in this way to avoid scrutiny.

OpenAI and Its Secrets?

The authors behind the lawsuit suspect OpenAI of willful copyright infringement. If it is proven that the company knew about the risks and used the datasets anyway, it could face enormous penalties. In cases of willful infringement, the court can increase damages to as much as $150,000 for each infringed work, which is approximately CZK 3,450,000. OpenAI argues that all reasons for deleting the datasets are protected by attorney-client privilege because in-house lawyers were involved in the decision. They even had a Slack channel originally called "excise-libgen," meaning something like "remove LibGen." One of the lawyers, Jason Kwon, suggested changing the name to "project-clear" to make it look less suspicious.

Judge Wang, however, reviewed the Slack messages and found that most of them were not privileged at all—they contained no legal advice, only ordinary discussions. According to her, OpenAI cannot designate an entire channel as confidential simply because a lawyer was mentioned in it. The authors hope these messages will reveal whether OpenAI deleted the datasets out of fear of legal trouble or whether it may still be using them under different names. One of the authors' lawyers, Christopher Young, suggested in a court filing that if it is proven that OpenAI decided not to use the datasets for newer models because of legal risks, it could seriously damage the company.

Judge Criticizes OpenAI for Misrepresentation

Wang was also angered by how OpenAI misrepresented a ruling in a separate case involving Anthropic. OpenAI claimed that Judge William Alsup had said downloading pirated books was legal if they were subsequently used to train AI. But that is nonsense—Alsup actually wrote that such piracy was "inherently, irredeemably infringing" and that no defendant could explain why it had not obtained the books legally. Wang called this a "bizarre" and "gross" misrepresentation and emphasized that OpenAI's conduct placed it squarely in the category Alsup had criticized: pirating data and then deleting it.

This dispute could affect the entire case. The authors also want to question Dario Amodei, the head of Anthropic, who allegedly created these datasets while he was still working at OpenAI. In March, the court ruled that Amodei must answer questions about his role in their deletion. OpenAI objected but lost. The company now says it disagrees with Judge Wang's ruling and plans to appeal.

Anthropic CEO Dario Amodei
Anthropic CEO Dario Amodei.

The Noose Is Tightening Around OpenAI

OpenAI tried to claim that it had acted in good faith, but then removed words such as "innocent" and "good faith" from court documents, which only strengthened the authors' suspicions. Wang noted that the jury has the right to know the basis for OpenAI's alleged good faith. This case is reminiscent of the recent settlement with Anthropic, in which authors received $1.5 billion, equivalent to about CZK 34.5 billion. Evidence emerged there that Anthropic had become less enthusiastic about training on pirated books for legal reasons. The authors hope to find similar evidence in OpenAI's messages.

The entire situation sheds light on how companies such as OpenAI collect data for AI. OpenAI claims that it did not misrepresent any information and merely used ambiguous wording that led to a misunderstanding. But Judge Wang rejected that argument and described its privilege claims as "far-fetched." OpenAI must now deal with the consequences or risk losing the entire case because of its own secrecy.

Source: arstechnica.com

Advertisement

Content created with help from UpTier.

SEO and GEO on autopilot. UpTier’s multi-agent systems write and optimize content for search engines and AI answers.

Discover UpTier ↗

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

OpenAI gives Codex reusable cloud workspaces accessible from any deviceOpenAI gives Codex reusable cloud workspaces accessible from any device
Codex gains reusable cloud development environments, alongside voice controls in its CLI, code reviews in the ChatGPT desktop app and cloud-based security tools.
2 min read
2. 10. 2026
Amazon releases Strands Decider 2B for AI workflow decisionsAmazon releases Strands Decider 2B for AI workflow decisions
Strands Decider 2B selects from predefined options and returns a confidence score. The fully open-source model is available now and small enough to run locally.
2 min read
1. 10. 2026
OpenAI says it disrupted a campaign to extract hidden model reasoningOpenAI says it disrupted a campaign to extract hidden model reasoning
OpenAI reported a coordinated effort to extract protected model reasoning and said it closed an extraction pathway. It attributed the main cluster of activity to individuals associated with Moonshot AI, the developer of Kimi.
3 min read
1. 10. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok