Unsealed NYT Case Filings Expose Internal Warnings at Microsoft and OpenAI

Microsoft maintains that its AI products are transformative fair use. The newly public material instead spotlights internal warnings that chatbots could drain traffic from the publishers whose work feeds them.

By 4 min read
Unsealed NYT Case Filings Expose Internal Warnings at Microsoft and OpenAI
Unsealed NYT Case Filings Expose Internal Warnings at Microsoft and OpenAI

Listen to this story

The audio brief

About 1:34
0:001:34
Read transcript
Newly unsealed court filings in The New York Times’ copyright case against Microsoft and OpenAI expose unusually direct internal warnings: AI answers could divert readers from the publishers whose reporting helps train the systems. One Microsoft memo from January 2023 called large-scale news scraping “the largest theft of labor in human history,” and warned that broad scraping could make a “complete mockery” of fair use. The plaintiffs also cite OpenAI’s Nick Turley describing commercial AI trained on news as an “existential threat” to publishers and “largely substitutive.” Microsoft chief executive Satya Nadella testified that chatbots can take clicks by answering on the AI platform instead of sending users to the source. The filing points to company records showing click-through declines ranging from 83 to 93 percent for some outlets, and 51 to 94 percent for others. It says Copilot produced 93 percent fewer referrals to The New York Times than traditional Bing searches—but that is evidence from the plaintiffs, not a court finding of causation. The publishers also allege Microsoft reused a Bing dataset, while OpenAI used 1.8 million Times articles despite commercial-use restrictions, with other allegations involving paywall evasion and removed copyright notices. Microsoft says those internal documents reflect one employee’s view and maintains that its products are transformative fair use. The open question is whether the alleged data path, reproduction risk, and referral losses together show competition with the news supply chain itself.

Story brief

3 key points

The filings put traffic substitution and data provenance at the center of the Microsoft–OpenAI copyright fight. Plaintiffs cite internal warnings that news-trained AI could undermine the publishers supplying its data, alongside alleged use of restricted Times material and plans to evade a paywall. Microsoft disputes the employees’ characterization and maintains that Copilot is transformative fair use. The evidence...

  1. 01

    Microsoft records cited by plaintiffs show news click-through declines of 83%–93% for some outlets and 51%–94% for others.

  2. 02

    The filing says Copilot reduced referrals to The New York Times by 93% versus traditional Bing searches; causation remains disputed.

  3. 03

    A January 2023 memo from Brent Hecht called large-scale news scraping “the largest theft of labor in human history.”

Microsoft says its AI products are transformative fair use and do not substitute for news sites. But newly unsealed filings in the New York Times-led copyright case put internal warnings beside that defense: company employees described a system in which AI answers could replace visits to publishers and erode the supply of reporting the models rely on.

The motion for summary judgment unsealed material from news plaintiffs’ case against Microsoft and OpenAI. It cites a January 2023 memo from Microsoft Director of Applied Science Brent Hecht calling large-scale news scraping “an astonishing theft of unprecedented proportions” and potentially “the largest theft of labor in human history.” The plaintiffs’ motion says Hecht also warned that broad scraping would make a “complete mockery” of fair use.

Internal accounts describe an answer engine as a substitute

According to the filing, OpenAI head of ChatGPT Nick Turley called commercial AI products trained on news an “existential threat” to publishers and said they were “largely substitutive.” Microsoft CEO Satya Nadella similarly testified that chatbots can take clicks by providing information on the AI platform rather than sending users to the underlying source.

It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its ‘content supply chain.’

Microsoft internal document, quoted in the news plaintiffs’ motion

The filing says Microsoft’s own records showed click-through-rate declines of 83% to 93% for some news plaintiffs and 51% to 94% for others. It also reports that Copilot reduced click-through to The New York Times by 93% compared with traditional Bing searches. Those figures are evidence cited by the plaintiffs, not a court finding that Copilot caused the declines.

Photo of Ashley Belanger
This story was updated with a quote from Steven Lieberman. Source: arstechnica.com.

The filing also alleges a disputed path from crawling to training

  • The plaintiffs allege Microsoft repurposed a Bing dataset for OpenAI training without consulting publishers.
  • They allege OpenAI used a third-party New York Times dataset containing 1.8 million articles despite commercial-use restrictions.
  • They also allege OpenAI employees planned to bypass the Times paywall without detection and that copyright notices were stripped from some training data.

Nadella testified that paywalled content should be licensed before it is used to train AI models. He said he would have supported retraining OpenAI models to exclude such content if he had known about the alleged scraping. Microsoft, however, said Hecht’s documents represented one employee’s perspective, not the company’s view or legal analysis, and defended its products as transformative fair use.

What the disclosures can and cannot settle

That limitation is especially important because the filing pairs striking internal language with contested legal conclusions. The publishers are using the documents to argue that Microsoft and OpenAI understood both sides of the alleged problem: training on news without permission and building products that reduce demand for the news itself. Microsoft rejects the premise that its AI products are substitutes.

The case now gives the fair-use fight a more concrete dispute than an abstract question about whether models learn from copyrighted work. It asks whether the alleged route to the data, the alleged ability to reproduce reporting, and the measured loss of referral traffic together show a product competing with its inputs. The court has yet to decide that question.

Editorial analysis

Our Read

The filing does not resolve whether training on news is fair use. But it sharpens a pressure point already central to the publisher cases: a model’s training use may be argued as transformative, while an answer product can still be accused of displacing the source it summarizes. The next consequential development is whether the court treats the alleged substitution evidence, including internal traffic data and executive testimony, as material to that distinction. The unsealed record also makes the route by which data was obtained harder to separate from the broader fair-use debate.

Sources

  1. techcrunch.comMicrosoft exec called AI scraping ‘the largest theft of labor in human history,' new unredacted filings reveal | TechCrunch
  2. arstechnica.comMicrosoft exec called AI scraping the “largest theft of labor in human history”
  3. news.ssbcrack.comNew revelations in The New York Times vs. OpenAI and Microsoft lawsuit highlight AI scraping as "theft" and threats to journalism. - SSBCrack News

Loading discussion...