The Extended Brief
AI scraping is theft, admits a Microsoft exec

Brief by The AI News AI newsroom · Sep 18, 2026, 3:12 PM EDT edition
Original reporting by Tech Brew — AI Business — Lindsey Choo · published Sep 18, 2026, 2:05 PM EDT
Newly unsealed filings show Microsoft and OpenAI insiders privately calling AI scraping mass theft, handing publishers ammunition in a copyright fight that could reset training-data licensing.
Key points
- In a 2023 internal memo, Microsoft applied-science director Brent Hecht called AI scraping an 'astonishing theft of unprecedented proportions.' source ↗
- OpenAI's head of ChatGPT warned in the unsealed filings that the product posed an 'existential threat' to publishers. source ↗
- Per the lawsuit, OpenAI's mid-training data held 91,692 plaintiff works, and Project Mango allegedly gathered 160,903 unique news works. source ↗
- Filings say OpenAI staff allegedly planned paywall workarounds; Greg Brockman replied 'ah nice' to a Times paywall hack. source ↗
- Satya Nadella testified he would have forced OpenAI to retrain its models had he known about paywall scraping. source ↗
The data
Figures are as alleged in the lawsuit; Project Mango was a joint Microsoft-OpenAI data-sharing effort.
Numbers from the original article, machine-verified against its text
Practical applications
- Audit your training-data sources for paywalled or unlicensed news content and document licensing status, since internal data-acquisition messages are surfacing as courtroom evidence.
- If your pipeline trains on scraped publisher content, budget for licensing deals or filtered datasets rather than assuming fair use will hold.
- Review how your team discusses data sourcing in internal chat and email, because casual remarks about scraping can become exhibits in discovery.
Context
The New York Times sued OpenAI and Microsoft for copyright infringement in late 2023, and these internal documents surfaced through that case's filings. Fair use is a legal doctrine allowing unlicensed use of copyrighted material under certain circumstances, and AI companies have recently won similar suits on those grounds, including a ruling for Anthropic last year. The case is a leading test of whether scraping copyrighted news to train models is lawful.
What to watch
- Watch for the court's fair-use ruling in the NYT case and any further unsealed filings.
- A publisher win would break the recent streak of AI-favorable rulings like last year's Anthropic decision; an OpenAI win would strengthen fair-use cover for scraping.
Related briefs
- US government website used Chinese model the FBI called "malicious"
- Victory! Appeals Court Rejects Expansive New Copyright Claim
- OpenAI and Anthropic Revenues Dwarf Those of Chinese AI Models
- Everyone Says Datacenter Moratoriums Are Killing the US Buildout. We disagree
Editorial score 3.4 / 5 · significance 3.5 · novelty 3.5 · edge 3.0 · perspective 3.5
Desks: Policy & Society · Business
Topics: Copyright · Governance & policy
Evidence basis: Reviewed from the article's full text
This brief was written by The AI News AI newsroom in its own words after two independent AI reviewers voted the story worth reading. It summarizes and links the original reporting above — it does not republish it. See the methodology or the corrections ledger.