New unredacted court filings in the New York Times’ copyright lawsuit against OpenAI and Microsoft allege large-scale scraping of paywalled news content for AI training, including internal executives’ characterizations of the conduct as theft
Legal action ● Confirmed 84% confidence first seen
Fresh unredacted filings submitted in the New York Times’ lawsuit against OpenAI and Microsoft allege that large-scale scraping of paywalled news articles was used for AI training. The documents also include internal statements from Microsoft and OpenAI leaders describing the activity in highly critical terms and contend it has caused major economic harm to publishers.
Decision brief
- What changed
- Newly unredacted filings in The New York Times’ copyright case against OpenAI and Microsoft allege that AI training used large-scale scraping of paywalled news content, including tens of thousands of copies of publisher works in training datasets. The filings also reveal internal statements from Microsoft and OpenAI leaders describing the scraping in terms such as theft and warning of harm to publishers and the web ecosystem.
- Why it matters
- For business leaders using or building on foundation models, this raises the legal and commercial significance of provenance, licensing, and indemnity in AI supply chains. The newly public internal characterizations may strengthen plaintiffs’ leverage in settlement, product restrictions, or future licensing demands, which could affect model costs, partner risk allocation, and enterprise adoption decisions.
- Evidence
- All three cited outlets report on the same unredacted court filings in the NYT lawsuit and consistently describe allegations of large-scale scraping of paywalled news content plus internal executive statements characterizing the conduct as theft. TechCrunch and Ars Technica both cite the filings’ descriptions of dataset contents and executive remarks, while 404 Media separately highlights the filings’ claims of economic harm to publishers, providing cross-outlet consistency based on the court record.
- What remains uncertain
- These are allegations and excerpts from court filings, not final judicial findings on liability, fair use, or damages. It remains unverified from the provided coverage how representative the cited datasets and internal remarks are across all model-training practices, and whether the case will materially change licensing terms, enterprise indemnities, or product availability.
- Monitor next
- Watch for the court’s next substantive ruling on the admissibility and significance of the unredacted evidence, especially any decision that narrows or expands exposure around training on publisher content.
Analytical support, not advice — assumptions and open questions stated above.