Anthropic’s $1.5B Deal Shows AI Training Fair Use Has a Data-Source Limit
Early U.S. decisions point to a narrower question than whether models can learn from copyrighted works: how the material was obtained, what it was authorized for, and whether the new product directly competes with it.
Listen to this story
The audio brief
Story brief
3 key pointsA $1.5 billion settlement involving Anthropic clarifies that the provenance of training data may matter as much as the training use itself. Judge William Alsup found that training on lawfully obtained books could be transformative and potentially fair, but did not extend that protection to a permanent corpus sourced from shadow libraries. Other disputes add different constraints: Ross lost a fair-use argument over a...
- 01
Alsup’s Bartz ruling separates model training from dataset retention: a fair-use finding for one does not immunize a pirated archive.
- 02
Thomson Reuters v. Ross suggests direct competition can weigh against fair use when AI training supports a substitute product.
- 03
A July 10 class action targets Google’s Gemini over alleged unlicensed use of books and journals from Google services.
Anthropic’s $1.5 billion settlement did not overturn a favorable court view of model training. It exposed a sharper limit: training on lawfully obtained books may qualify as fair use, while maintaining a library assembled from pirated copies can bring separate copyright exposure.
The split inside Anthropic’s case
Fair use is a copyright exception assessed through factors including a use’s purpose and nature, the amount used, and market effects. A key issue is whether the new use is transformative, meaning it has a further purpose or different character.
In Bartz v. Anthropic, Judge William Alsup found training large language models on lawfully obtained books transformative and potentially fair use. He characterized the activity as closer to reading and studying literature than copying it. The finding did not cover Anthropic’s permanent library of books from shadow libraries, including Library Genesis and Pirate Library Mirror.
A competing product produced the opposite result
The result was different in Thomson Reuters v. Ross Intelligence. Judge Stephanos Bibas ruled that Ross’s use of Thomson Reuters content to build a competing AI legal platform was not fair use because it was not transformative. The contrast leaves direct competition as a consequential part of the legal analysis, not merely the existence of copyrighted training material.
Google faces a permissions-based challenge
Hachette, Cengage, Elsevier, Scott Turow, and S.C.R.I.B.E. filed a class action against Google on July 10, alleging that Gemini was trained on copyrighted books and journal articles without the necessary rights. The complaint says works obtained through Google Books, Google Play Books, and other Google services were used without fresh permission from rights holders.
It also alleges Google removed or altered copyright-management information, a theory that could create separate Digital Millennium Copyright Act claims. Those allegations are unproven, and Google can respond. The case therefore turns on claimed permissions and handling of data, rather than the shadow-library conduct at issue for Anthropic.
Inputs and outputs remain different legal questions
The law governing training data is separate from whether AI output can be copyrighted. In Thaler v. Perlmutter, a court ruled that a work generated entirely by AI is not copyrightable.
U.S. copyright law has not been updated since 1976, and most AI companies remain in pending litigation over training. The current cases are influential but not final: they offer boundaries around acquisition, permission, and competition, not a universal answer for every dataset.
Sources
- startupfortune.comHow One Judge's Split Ruling on Anthropic Became AI's Copyright Rulebook - Startup Fortune
- techcrunch.comIs it legal to train AI models on copyrighted books? It’s complicated | TechCrunch