Unsealed NYT Case Filings Expose Internal Warnings at Microsoft and OpenAI
Microsoft maintains that its AI products are transformative fair use. The newly public material instead spotlights internal warnings that chatbots could drain traffic from the publishers whose work feeds them.
Listen to this story
The audio brief
Story brief
3 key pointsThe filings put traffic substitution and data provenance at the center of the Microsoft–OpenAI copyright fight. Plaintiffs cite internal warnings that news-trained AI could undermine the publishers supplying its data, alongside alleged use of restricted Times material and plans to evade a paywall. Microsoft disputes the employees’ characterization and maintains that Copilot is transformative fair use. The evidence...
- 01
Microsoft records cited by plaintiffs show news click-through declines of 83%–93% for some outlets and 51%–94% for others.
- 02
The filing says Copilot reduced referrals to The New York Times by 93% versus traditional Bing searches; causation remains disputed.
- 03
A January 2023 memo from Brent Hecht called large-scale news scraping “the largest theft of labor in human history.”
Microsoft says its AI products are transformative fair use and do not substitute for news sites. But newly unsealed filings in the New York Times-led copyright case put internal warnings beside that defense: company employees described a system in which AI answers could replace visits to publishers and erode the supply of reporting the models rely on.
The motion for summary judgment unsealed material from news plaintiffs’ case against Microsoft and OpenAI. It cites a January 2023 memo from Microsoft Director of Applied Science Brent Hecht calling large-scale news scraping “an astonishing theft of unprecedented proportions” and potentially “the largest theft of labor in human history.” The plaintiffs’ motion says Hecht also warned that broad scraping would make a “complete mockery” of fair use.
Internal accounts describe an answer engine as a substitute
According to the filing, OpenAI head of ChatGPT Nick Turley called commercial AI products trained on news an “existential threat” to publishers and said they were “largely substitutive.” Microsoft CEO Satya Nadella similarly testified that chatbots can take clicks by providing information on the AI platform rather than sending users to the underlying source.
It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its ‘content supply chain.’
Microsoft internal document, quoted in the news plaintiffs’ motion
The filing says Microsoft’s own records showed click-through-rate declines of 83% to 93% for some news plaintiffs and 51% to 94% for others. It also reports that Copilot reduced click-through to The New York Times by 93% compared with traditional Bing searches. Those figures are evidence cited by the plaintiffs, not a court finding that Copilot caused the declines.
The filing also alleges a disputed path from crawling to training
- The plaintiffs allege Microsoft repurposed a Bing dataset for OpenAI training without consulting publishers.
- They allege OpenAI used a third-party New York Times dataset containing 1.8 million articles despite commercial-use restrictions.
- They also allege OpenAI employees planned to bypass the Times paywall without detection and that copyright notices were stripped from some training data.
Nadella testified that paywalled content should be licensed before it is used to train AI models. He said he would have supported retraining OpenAI models to exclude such content if he had known about the alleged scraping. Microsoft, however, said Hecht’s documents represented one employee’s perspective, not the company’s view or legal analysis, and defended its products as transformative fair use.
What the disclosures can and cannot settle
That limitation is especially important because the filing pairs striking internal language with contested legal conclusions. The publishers are using the documents to argue that Microsoft and OpenAI understood both sides of the alleged problem: training on news without permission and building products that reduce demand for the news itself. Microsoft rejects the premise that its AI products are substitutes.
The case now gives the fair-use fight a more concrete dispute than an abstract question about whether models learn from copyrighted work. It asks whether the alleged route to the data, the alleged ability to reproduce reporting, and the measured loss of referral traffic together show a product competing with its inputs. The court has yet to decide that question.
Editorial analysis
Our Read
The filing does not resolve whether training on news is fair use. But it sharpens a pressure point already central to the publisher cases: a model’s training use may be argued as transformative, while an answer product can still be accused of displacing the source it summarizes. The next consequential development is whether the court treats the alleged substitution evidence, including internal traffic data and executive testimony, as material to that distinction. The unsealed record also makes the route by which data was obtained harder to separate from the broader fair-use debate.
Sources
- techcrunch.comMicrosoft exec called AI scraping ‘the largest theft of labor in human history,' new unredacted filings reveal | TechCrunch
- arstechnica.comMicrosoft exec called AI scraping the “largest theft of labor in human history”
- news.ssbcrack.comNew revelations in The New York Times vs. OpenAI and Microsoft lawsuit highlight AI scraping as "theft" and threats to journalism. - SSBCrack News
Loading discussion...
Reader comments
Newest comments first. Replies stay oldest first.