OpenAI has found itself in serious trouble. At the center of attention is a class-action lawsuit filed by book authors who claim that their works were illegally used to train artificial intelligence. The key issue? OpenAI deleted two datasets full of pirated books, called Books1 and Books2, before launching ChatGPT in 2022. These datasets were created by former OpenAI employees in 2021 by crawling the open web and stealing most of the data from a so-called shadow library known as Library Genesis, or LibGen for short. OpenAI claims that it stopped using the datasets that same year and decided internally to delete them. But the authors believe there is more to the story, and it now appears that the court may agree with them.
Last week, U.S. Magistrate Judge Ona Wang ruled that OpenAI must share all internal communications with the company's lawyers regarding the deletion of these datasets. It must also provide all internal references to LibGen that OpenAI had withheld under the guise of attorney-client privilege. Wang ordered OpenAI to turn over these materials by December 8 and make its in-house lawyers available for questioning by December 19. Why such a ruling? According to the judge, OpenAI acted improperly—first claiming that the datasets being "unused" was the reason for deleting them, then withdrawing that claim and suddenly declaring all reasons privileged. Wang described this as "vacillation" and said that OpenAI cannot shift its positions in this way to avoid scrutiny.
OpenAI and Its Secrets?
The authors behind the lawsuit suspect OpenAI of willful copyright infringement. If it is proven that the company knew about the risks and used the datasets anyway, it could face enormous penalties. In cases of willful infringement, the court can increase damages to as much as $150,000 for each infringed work, which is approximately CZK 3,450,000. OpenAI argues that all reasons for deleting the datasets are protected by attorney-client privilege because in-house lawyers were involved in the decision. They even had a Slack channel originally called "excise-libgen," meaning something like "remove LibGen." One of the lawyers, Jason Kwon, suggested changing the name to "project-clear" to make it look less suspicious.
Judge Wang, however, reviewed the Slack messages and found that most of them were not privileged at all—they contained no legal advice, only ordinary discussions. According to her, OpenAI cannot designate an entire channel as confidential simply because a lawyer was mentioned in it. The authors hope these messages will reveal whether OpenAI deleted the datasets out of fear of legal trouble or whether it may still be using them under different names. One of the authors' lawyers, Christopher Young, suggested in a court filing that if it is proven that OpenAI decided not to use the datasets for newer models because of legal risks, it could seriously damage the company.
Judge Criticizes OpenAI for Misrepresentation
Wang was also angered by how OpenAI misrepresented a ruling in a separate case involving Anthropic. OpenAI claimed that Judge William Alsup had said downloading pirated books was legal if they were subsequently used to train AI. But that is nonsense—Alsup actually wrote that such piracy was "inherently, irredeemably infringing" and that no defendant could explain why it had not obtained the books legally. Wang called this a "bizarre" and "gross" misrepresentation and emphasized that OpenAI's conduct placed it squarely in the category Alsup had criticized: pirating data and then deleting it.
This dispute could affect the entire case. The authors also want to question Dario Amodei, the head of Anthropic, who allegedly created these datasets while he was still working at OpenAI. In March, the court ruled that Amodei must answer questions about his role in their deletion. OpenAI objected but lost. The company now says it disagrees with Judge Wang's ruling and plans to appeal.
The Noose Is Tightening Around OpenAI
OpenAI tried to claim that it had acted in good faith, but then removed words such as "innocent" and "good faith" from court documents, which only strengthened the authors' suspicions. Wang noted that the jury has the right to know the basis for OpenAI's alleged good faith. This case is reminiscent of the recent settlement with Anthropic, in which authors received $1.5 billion, equivalent to about CZK 34.5 billion. Evidence emerged there that Anthropic had become less enthusiastic about training on pirated books for legal reasons. The authors hope to find similar evidence in OpenAI's messages.
The entire situation sheds light on how companies such as OpenAI collect data for AI. OpenAI claims that it did not misrepresent any information and merely used ambiguous wording that led to a misunderstanding. But Judge Wang rejected that argument and described its privilege claims as "far-fetched." OpenAI must now deal with the consequences or risk losing the entire case because of its own secrecy.
Source: arstechnica.com



