AI Company May Train on Legally Obtained Books, but Piracy Is Prohibited, U.S. Court Rules
A U.S. federal court in California has issued an important ruling in a case brought by writers Andrea Bartz, Charles Graeber, and Kirk Wallace Johnson against Anthropic. The judge ruled that training artificial intelligence on legally purchased books falls under the concept of fair use, while rejecting any defense concerning the millions of pirated copies in the company's digital library.
Spectacularly Transformative Use
The judge described AI training as a "spectacularly" transformative process. According to the ruling, the Claude language model can be compared to aspiring writers who learn from established authors but do not copy them. This comparison was crucial to the ruling that the use of legally obtained books was permissible. The authors failed to prove in their lawsuit that Claude was capable of generating outputs similar to their original works. This significantly weakened their main arguments about the competitive harm caused by the AI system.
Millions Spent on Legal Books Versus Pirated Downloads
Court filings revealed that Anthropic legally spent "many millions" purchasing printed books. These books were then scanned into digital files for use in training artificial intelligence. The court considered this process legitimate. On the other hand, the company also downloaded millions of books from piracy websites and permanently stored them in its central library. The goal was to create a library containing "all the books in the world" and preserve them "forever." The court found that this conduct infringed the authors' copyrights.
Further Proceedings and Implications
Anthropic will face trial in December over the willful infringement of copyrights in pirated works. Potential damages could reach up to $150,000 per book, posing a significant financial risk to the company.
This ruling gives AI laboratories the green light to train on legally obtained data. It could become one of the first precedents for the countless cases currently being brought against technology companies. Given the lack of clarity about how much copyrighted material is included in training data, this is only one battle in what appears to be a very long legal war between AI companies and copyright holders.



