Anthropic’s $1.5B Deal Shows AI Training Fair Use Has a Data-Source Limit

Early U.S. decisions point to a narrower question than whether models can learn from copyrighted works: how the material was obtained, what it was authorized for, and whether the new product directly competes with it.

By 2 min read
Anthropic’s $1.5B Deal Shows AI Training Fair Use Has a Data-Source Limit
Anthropic’s $1.5B Deal Shows AI Training Fair Use Has a Data-Source Limit

Listen to this story

The audio brief

About 1:41
0:001:41
Read transcript
Anthropic’s one-point-five-billion-dollar settlement has exposed a sharper limit on AI training and fair use. A judge found that training a model on lawfully obtained books could be transformative and potentially fair. But that protection did not extend to Anthropic’s permanent library of books collected from shadow libraries, including Library Genesis and Pirate Library Mirror. The distinction matters because fair use is not a blanket defense for every step in a training pipeline. Courts weigh the purpose of the use, how much material was taken, and the effect on the market. In Bartz versus Anthropic, Judge William Alsup compared model training with reading and studying literature. In Thomson Reuters versus Ross Intelligence, the analysis went the other way: Judge Stephanos Bibas found that Ross’s use of Thomson Reuters content to build a competing legal platform was not transformative. Direct competition, in other words, can be decisive. Google now faces a different challenge. Publishers Hachette, Cengage, and Elsevier, along with Scott Turow and S.C.R.I.B.E., allege that Gemini used books and journal articles from Google Books, Google Play Books, and other services without the necessary permissions. They also allege altered copyright-management information, which could support separate Digital Millennium Copyright Act claims. Those allegations remain unproven. And training data is a separate question from output: Thaler versus Perlmutter held that wholly AI-generated work is not copyrightable. The boundary to watch is how courts treat acquisition, authorization, and competition together—not simply whether copyrighted works entered training.

Story brief

3 key points

A $1.5 billion settlement involving Anthropic clarifies that the provenance of training data may matter as much as the training use itself. Judge William Alsup found that training on lawfully obtained books could be transformative and potentially fair, but did not extend that protection to a permanent corpus sourced from shadow libraries. Other disputes add different constraints: Ross lost a fair-use argument over a...

  1. 01

    Alsup’s Bartz ruling separates model training from dataset retention: a fair-use finding for one does not immunize a pirated archive.

  2. 02

    Thomson Reuters v. Ross suggests direct competition can weigh against fair use when AI training supports a substitute product.

  3. 03

    A July 10 class action targets Google’s Gemini over alleged unlicensed use of books and journals from Google services.

Anthropic’s $1.5 billion settlement did not overturn a favorable court view of model training. It exposed a sharper limit: training on lawfully obtained books may qualify as fair use, while maintaining a library assembled from pirated copies can bring separate copyright exposure.

The split inside Anthropic’s case

Fair use is a copyright exception assessed through factors including a use’s purpose and nature, the amount used, and market effects. A key issue is whether the new use is transformative, meaning it has a further purpose or different character.

In Bartz v. Anthropic, Judge William Alsup found training large language models on lawfully obtained books transformative and potentially fair use. He characterized the activity as closer to reading and studying literature than copying it. The finding did not cover Anthropic’s permanent library of books from shadow libraries, including Library Genesis and Pirate Library Mirror.

A competing product produced the opposite result

The result was different in Thomson Reuters v. Ross Intelligence. Judge Stephanos Bibas ruled that Ross’s use of Thomson Reuters content to build a competing AI legal platform was not fair use because it was not transformative. The contrast leaves direct competition as a consequential part of the legal analysis, not merely the existence of copyrighted training material.

Google faces a permissions-based challenge

Hachette, Cengage, Elsevier, Scott Turow, and S.C.R.I.B.E. filed a class action against Google on July 10, alleging that Gemini was trained on copyrighted books and journal articles without the necessary rights. The complaint says works obtained through Google Books, Google Play Books, and other Google services were used without fresh permission from rights holders.

It also alleges Google removed or altered copyright-management information, a theory that could create separate Digital Millennium Copyright Act claims. Those allegations are unproven, and Google can respond. The case therefore turns on claimed permissions and handling of data, rather than the shadow-library conduct at issue for Anthropic.

Inputs and outputs remain different legal questions

The law governing training data is separate from whether AI output can be copyrighted. In Thaler v. Perlmutter, a court ruled that a work generated entirely by AI is not copyrightable.

U.S. copyright law has not been updated since 1976, and most AI companies remain in pending litigation over training. The current cases are influential but not final: they offer boundaries around acquisition, permission, and competition, not a universal answer for every dataset.

Sources

  1. startupfortune.comHow One Judge's Split Ruling on Anthropic Became AI's Copyright Rulebook - Startup Fortune
  2. techcrunch.comIs it legal to train AI models on copyrighted books? It’s complicated | TechCrunch

Loading discussion...