Microsoft's Own Documents Call AI Scraping 'Theft' While Its Lawyers Argue Fair Use
Unsealed filings from the NYT v. OpenAI/Microsoft copyright case contain internal memos, a 93% traffic-drop figure, and a paywall-bypass reply that the defendant-side legal team will have to explain.
What happened
New unredacted documents from the NYT v. OpenAI and Microsoft copyright lawsuit went public Thursday, drawn from a motion for summary judgment filed by the news plaintiffs. Internal communications in the filings describe AI scraping as "theft" and concede the companies' products directly substitute for publisher traffic.
Context
The New York Times sued OpenAI and Microsoft roughly three years ago, alleging both companies trained generative AI models on its copyrighted content without permission. Judges have been described as largely favorable to the AI companies' fair-use defence; in August 2026 the Trump administration filed a brief supporting OpenAI's unlicensed use of copyrighted material for LLM training. What surfaced Thursday comes primarily from the Times' own brief; underlying exhibits remain sealed, and the quoted passages lack their original context.
How it works
OpenAI built training corpora called WebText and WebText2 that disproportionately relied on scraped news content, drawing from Common Crawl and the Bing Index. One dataset held over 2 million documents from nytimes.com alone; mid-training sets contained 91,692+ copies of works from the NYT, Daily News, and Center for Investigative Reporting. Microsoft received the full GPT-3 training dataset from OpenAI and fed data back through projects called Taxi and Mango.
The filings allege deliberate removal of copyright notices from training data, the stated reason being that researchers "wouldn't want model outputting" notices to users. A researcher told Greg Brockman about a technique to bypass the nytimes.com paywall; Brockman replied "ah nice."
The figure that matters most: Microsoft's own data shows its Copilot answer engine cut NYT-domain click-through rates by up to 93% versus traditional Bing search. Brent Hecht, a Microsoft Director of Applied Science, called the resulting publisher-traffic decline a "doom loop" in a January 2024 internal presentation.
Our read
The most interesting thing in these filings is not any single quote. It is the gap between what Microsoft's lawyers tell the court and what its own scientists wrote privately. Brent Hecht called AI scraping "the largest theft of labor in human history" in a January 2023 memo. Ars Technica reports he hedged with "perhaps" before that phrase; TechCrunch's version omits the word. Either way, a defendant-side executive used "theft" in a private document while the company's public position is that the same activity is fair use.
More consequential is the 93% figure. Fair-use doctrine's fourth prong asks whether the use harms the market for the original work. Microsoft's own data says its product eliminated 93% of the traffic that would have reached the NYT. Nick Turley, OpenAI's head of ChatGPT, called the output "largely substitutive." Brockman called the models "excellent at news." That is not a neutral description of a tool that supplements journalism. It is an internal admission that the tool replaces it.
The limitation: these quotes come from the NYT's brief, not from unsealed exhibits. Context is unavailable, neither company returned TechCrunch's comment requests, and the court has not ruled. This is one party's curated selection from a much larger record.
What this changes
For a studio running local diffusion models, the direct operational change is near-zero. These filings do not alter the licence terms of the open-weight models you already load into ComfyUI, and you are consuming pre-trained weights, not training from raw scraped corpora.
Three cautions still apply. If you fine-tune or train LoRAs on images or text scraped without a licence, the substitution and market-harm arguments developed here apply to you in miniature. If your generated video directly replaces licensed stock footage in a client deliverable, the "doom loop" framing in these documents is the exact argument a rights-holder will use. And if you re-encode third-party assets into your pipeline, retaining provenance metadata is legal hygiene; the filings treat its removal as an aggravating factor.
License
No licence applies. This is a copyright-litigation story, not a model release or software distribution.
Key takeaways
- A Microsoft Director of Applied Science called AI scraping "the largest theft of labor in human history" in a 2023 internal memo, contradicting the company's public fair-use position.
- Microsoft's own data shows Copilot cut NYT click-through rates by up to 93%, the strongest single point against the fair-use "market harm" prong.
- The quotes come from the NYT's brief, not unsealed exhibits; neither company commented, and no ruling has been issued.
- For a studio consuming pre-trained open-weight models, the operational impact is minimal; the risk surface is in fine-tuning on unlicensed scraped data or generating output that directly substitutes for licensed content.
- The alleged removal of copyright metadata from training data is framed as an aggravating factor; retaining provenance in your own pipelines is the corresponding defensive step.
Sources
How this post was made
Drafted from clustered primary sources by the models below, then read, edited and approved by a human before it was published. The sources are listed in full at the end of the article.
- Drafted
- Independent sources
- 2
- cluster pair
- gemma4:12b
- cluster label
- gemma4:12b
- radar brief
- gemma4:12b
- research brief
- qwen3.8:27b
- draft article
- qwen3.8:27b
- short script
- qwen3.8:27b
- seo pack
- gemma4:12b
- Run
- editorial-20260917T223326Z