Hacker News front pageRebecca Bellan4 min readintermediate
Microsoft exec called AI scraping 'the largest theft of labor in human history'
Summary
New unredacted filings in the NYT v. OpenAI/Microsoft case show internal admissions that the companies scraped paywalled news content at massive scale, stripped copyright notices, and acknowledge that their products cannibalize traffic, raising serious fair‑use and copyright concerns.
- Internal Microsoft and OpenAI documents describe scraping >2 M NYT pages via Common Crawl and >91 k copies of works from major news outlets.
- Executives called the practice “theft” and warned it creates a “doom loop” that hurts both publishers and model performance.
- Microsoft’s own data shows Copilot can cut NYT click‑through rates by up to 93 % versus traditional Bing search.
- The filings allege deliberate paywall circumvention, stripping of copyright notices, and reciprocal data‑exchange programs (Project Taxi/Mango).
If the allegations are accurate, the scale of unlicensed data ingestion could set a precedent for how LLM developers source training data, potentially prompting stricter licensing regimes, new compliance tooling, and shifts in how publishers protect content.
4/10



