Microsoft exec called AI scraping ‘the largest theft of labor in human history,’ new unredacted filings reveal
TechCrunch Rebecca Bellan ● Covered by 3 sources
Microsoft and OpenAI’s own filings now call AI scraping theft. The surprise is how bluntly they say chatbot training can hurt publishers and the people who wrote the work.
Based on reporting by TechCrunch, Rebecca Bellan — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
A new batch of unredacted filing material in The New York Times’ lawsuit against OpenAI and Microsoft reads less like a legal defense than an accidental confession. In private, a Microsoft executive described the companies’ AI training practices as “theft.” OpenAI leadership, meanwhile, warned that its own products could be an “existential threat” to the publishers and journalists whose work helped train them.
The material is showing up now because of the Times’ filing, not because the underlying exhibits were opened up. Still, the quotes paint a messy picture. The companies are accused of building training sets by mass scraping, sneaking past paywalls, and stripping copyright notices from data before it reached the models. The case itself is three years old, and it still sits at the center of the bigger fight over whether training generative AI on copyrighted work is lawful under fair use.
That fair-use argument has been the backbone of the AI companies’ defense, and judges have often been receptive to it. But several of these admissions cut in the opposite direction. Microsoft’s own data says Copilot sent click-through rates for The New York Times’ domain down by as much as 93% compared with traditional Bing search. Satya Nadella also testified that anything behind a paywall should be licensed for grounding or training, and that if he had known OpenAI trained on paywalled material, he would have pushed for retraining.
OpenAI’s internal language is just as stark. Nick Turley, who leads ChatGPT, called publishers’ position an “existential threat,” and said the tools are “largely substitutive” and likely to become more so. Greg Brockman called the models “excellent at news.” Nadella agreed that chatting with a bot can replace going to the original source on the website. That is a problem for the fair-use story, because substitution is exactly what the doctrine is supposed to avoid.
The scale alleged in the filing is hard to shrug off. OpenAI’s mid-training datasets alone reportedly include more than 91,692 copies of works from The New York Times, Daily News, and Center for Investigative Reporting. A Common Crawl-based dataset included more than 2 million documents from nytimes.com alone. One Microsoft memo called the situation “an astonishing theft of unprecedented proportions” and even “the largest theft of labor in human history.”
My take — AI-written commentary, not fact-checked reporting
The industry has spent years pretending “training data” is a magic phrase that cancels out copyright, wages, and consent. It doesn’t. If a model is built from paywalled reporting and then used to siphon readers away from the source, that’s not innovation — that’s a very expensive way to be rude.
Read more about this at: TechCrunch