Policypublished

Anthropic’s $1.5B Deal Shows AI Training Fair Use Has a Data-Source Limit

Early U.S. decisions point to a narrower question than whether models can learn from copyrighted works: how the material was obtained, what it was authorized for, and whether the new product directly competes with it.

By 2 min read
Anthropic’s $1.5B Deal Shows AI Training Fair Use Has a Data-Source Limit

Listen to this story

The audio brief

About 1:32
0:001:32
Read transcript
Anthropic agreed to pay one point five billion dollars in a settlement that may narrow, rather than settle, the fight over AI training data. In Bartz versus Anthropic, Judge William Alsup found that training a language model on lawfully obtained books could be transformative and potentially fair use—closer to reading and studying than to republishing. But that protection did not extend to Anthropic’s permanent archive of books taken from shadow libraries, including Library Genesis and the Pirate Library Mirror. The distinction is crucial: a court may treat the act of training differently from the way a company acquired and retained its dataset. Another case points to a second constraint. In Thomson Reuters versus Ross Intelligence, Judge Stephanos Bibas rejected Ross’s fair-use defense, finding that its use of Thomson Reuters material supported a competing AI legal product. Direct market competition, in other words, can matter alongside the source of the data. Google now faces a permissions-based challenge. Publishers including Hachette, Cengage, and Elsevier, along with Scott Turow and S.C.R.I.B.E., allege that Gemini used books and journals without the necessary rights, and that copyright-management information was altered. Those claims remain unproven. And training is separate from output: Thaler versus Perlmutter held that wholly AI-generated work is not copyrightable. The constraint to watch is whether future rulings focus less on AI training in the abstract, and more on acquisition, permission, retention, and competition.

Story brief

3 key points

A $1.5 billion settlement involving Anthropic clarifies that the provenance of training data may matter as much as the training use itself. Judge William Alsup found that training on lawfully obtained books could be transformative and potentially fair, but did not extend that protection to a permanent corpus sourced from shadow libraries. Other disputes add different constraints: Ross lost a fair-use argument over a...

  1. 01

    Alsup’s Bartz ruling separates model training from dataset retention: a fair-use finding for one does not immunize a pirated archive.

  2. 02

    Thomson Reuters v. Ross suggests direct competition can weigh against fair use when AI training supports a substitute product.

  3. 03

    A July 10 class action targets Google’s Gemini over alleged unlicensed use of books and journals from Google services.

Anthropic’s $1.5 billion settlement did not overturn a favorable court view of model training. It exposed a sharper limit: training on lawfully obtained books may qualify as fair use, while maintaining a library assembled from pirated copies can bring separate copyright exposure.

The split inside Anthropic’s case

Fair use is a copyright exception assessed through factors including a use’s purpose and nature, the amount used, and market effects. A key issue is whether the new use is transformative, meaning it has a further purpose or different character.

In Bartz v. Anthropic, Judge William Alsup found training large language models on lawfully obtained books transformative and potentially fair use. He characterized the activity as closer to reading and studying literature than copying it. The finding did not cover Anthropic’s permanent library of books from shadow libraries, including Library Genesis and Pirate Library Mirror.

A competing product produced the opposite result

The result was different in Thomson Reuters v. Ross Intelligence. Judge Stephanos Bibas ruled that Ross’s use of Thomson Reuters content to build a competing AI legal platform was not fair use because it was not transformative. The contrast leaves direct competition as a consequential part of the legal analysis, not merely the existence of copyrighted training material.

Google faces a permissions-based challenge

Hachette, Cengage, Elsevier, Scott Turow, and S.C.R.I.B.E. filed a class action against Google on July 10, alleging that Gemini was trained on copyrighted books and journal articles without the necessary rights. The complaint says works obtained through Google Books, Google Play Books, and other Google services were used without fresh permission from rights holders.

It also alleges Google removed or altered copyright-management information, a theory that could create separate Digital Millennium Copyright Act claims. Those allegations are unproven, and Google can respond. The case therefore turns on claimed permissions and handling of data, rather than the shadow-library conduct at issue for Anthropic.

Inputs and outputs remain different legal questions

The law governing training data is separate from whether AI output can be copyrighted. In Thaler v. Perlmutter, a court ruled that a work generated entirely by AI is not copyrightable.

U.S. copyright law has not been updated since 1976, and most AI companies remain in pending litigation over training. The current cases are influential but not final: they offer boundaries around acquisition, permission, and competition, not a universal answer for every dataset.

Sources

  1. startupfortune.comHow One Judge's Split Ruling on Anthropic Became AI's Copyright Rulebook - Startup Fortune
  2. techcrunch.comIs it legal to train AI models on copyrighted books? It’s complicated | TechCrunch