← All articles
Sep 20, 2026

Microsoft's Own Memo Called AI Training Data Scraping 'The Largest Theft in History'

Unsealed court filings show a Microsoft director privately calling AI scraping the largest theft of labor in history. Here is what the documents reveal.

A newly unsealed memo from inside Microsoft has put a blunt phrase into the public record of one of the biggest lawsuits in AI. In January 2023, Microsoft director of applied science Brent Hecht wrote internally that AI training data scraping amounted to the largest theft of labor in human history. That memo only became public this week, more than three years after it was written, because a discovery process in The New York Times' copyright suit against OpenAI and Microsoft forced it into the open. The lawsuit itself has run for three years without producing a moment quite like this one: a senior person inside one of the two defendant companies describing, in writing, exactly what critics of AI training data collection have argued for years.

The filings were unsealed on September 17, and TechCrunch's reporting off the documents is the fullest account so far of what they contain. For anyone building products on top of frontier models, the interesting part isn't just the memo's language. It's the gap it reveals between what people inside these companies say to each other and what they say in public.

What the unsealed filings actually show

Hecht went further than that single line. In a separate internal Microsoft presentation a year later, in January 2024, he wrote that it is highly unusual for an end product to threaten the economic foundations of its essential suppliers, and said plainly that this was the situation Microsoft had created for its own large language model business with respect to its content supply chain. That is effectively an admission that the AI business was built on top of a content supply chain it was actively damaging.

He wasn't alone in saying more privately than publicly. Nick Turley, who leads ChatGPT at OpenAI, is quoted describing the threat to publishers as existential, telling colleagues that AI products are already substitutive for reading the original source and will keep getting more substitutive as they improve. And Microsoft CEO Satya Nadella, testifying under oath, said that anything paywalled should be licensed by whoever wants to use it for grounding or training a model, and that if he had known OpenAI was scraping paywalled content he would have required the models to be retrained.

None of that is standard public messaging from either company. It's the kind of thing that gets said in a deposition or an internal memo, not a blog post.

The scale behind the openai new york times lawsuit

The numbers in the filings are what turn this from a quote into a documented pattern. OpenAI's own training datasets reportedly contained more than 91,692 copies of content from the Times, the Daily News, and the Center for Investigative Reporting. A separate common crawl training data pull is said to have collected more than 2 million documents from nytimes.com alone. Two internally named efforts, Project Taxi and Project Mango, are described in the filings as the mechanism for assembling this material, with Project Mango alone tied to more than 160,903 unique works from news publishers.

The filings also describe practices around how that content was collected, including stripping copyright notices and working around paywalls rather than treating them as a boundary. That detail matters for anyone thinking about their own data pipeline: paywalls and copyright notices are commonly read as clear signals of restricted use, and the allegation here is that they were engineered around rather than respected.

There's also a real-world number attached to the substitution argument Turley made privately. According to the filings, click-through rates to the Times' own domain fell by roughly 93 percent when readers got information through Microsoft's Copilot compared to traditional Bing search results. That's the same substitution effect Nadella described in his testimony, measured directly.

Why this matters beyond one lawsuit

This case has been running for three years, and it hasn't been a candidate for coverage here until now, because most of what surfaced in earlier rounds was procedural. What changed is the specificity. A private memo calling something theft is one thing. A private memo calling it theft, backed by exact document counts, named internal projects, and a measured drop in referral traffic, is a documented pattern rather than a rhetorical flourish.

For anyone building AI products, the practical read isn't really about this one lawsuit. It's about what training data provenance disputes are turning into: a real, ongoing legal and reputational risk rather than a background concern. Vendors are increasingly signing licensing deals with publishers precisely because the alternative, arguing that broad scraping was fair use all along, gets harder to defend once a company's own executives are on record calling it something else. If you're evaluating a model provider, choosing a dataset, or building a RAG pipeline on scraped web content, the questions worth asking now are concrete ones: where did this training or grounding data actually come from, was it collected with the rights holder's knowledge, and what happens to your product if that provenance gets challenged later.

The gap between what a company says privately and what it says publicly isn't unique to AI. But this case is a clean example of how long that gap can survive before a lawsuit forces it into daylight, and how much more damaging the private version sounds once it's attached to real numbers.

Source: TechCrunch's reporting on the unsealed filings.

Join the newsletter

AI workflows and systems, straight to your inbox.

No spam. Unsubscribe anytime.