A Microsoft executive privately described AI training practices as ‘an astonishing theft of unprecedented proportions’ — and that’s just one of several admissions now surfacing in the three-year-old copyright lawsuit between The New York Times, OpenAI, and Microsoft. According to TechCrunch, newly unredacted court filings reveal a stark gap between what these companies said publicly and what their own executives wrote internally.
The lawsuit, originally filed by The New York Times in 2023, alleged that OpenAI and Microsoft violated copyright law by training generative AI models on its content without permission or payment. The central legal question, whether AI training qualifies as ‘fair use,’ still has no definitive answer. Courts have generally leaned toward AI companies on this point, and earlier this year the Trump administration filed a brief supporting OpenAI’s position. But the newly unsealed material cuts directly against that defense.
Fair use has four factors, and one of the most important is whether the use harms the market for the original work. Microsoft’s own internal data shows its Copilot ‘answer engine’ caused click-through rates to NYT’s domain to drop by as much as 93% compared to traditional Bing search. Brent Hecht, Microsoft’s Director of Applied Science, described this in a January 2024 internal presentation as a ‘doom loop’ that would damage both model performance and the broader web. The document also warned that it is ‘highly unusual that an end-product threatens the economic foundations of its essential suppliers’ — and yet that is exactly what the company acknowledged it had created.
OpenAI’s side of the record is equally damaging. Nick Turley, Head of ChatGPT, wrote internally that publishers face an ‘existential threat’ from the chatbot because it is ‘largely substitutive’ and will become more so over time. Greg Brockman described the models as ‘excellent at news.’ CEO Satya Nadella testified under oath that paywalled content should require a license and said he would have required OpenAI to retrain its models had he known they scraped behind-paywall material.
The scale of copying described in the filings is significant. OpenAI’s mid-training datasets reportedly contain more than 91,692 copies of works from the NYT, Daily News, and Center for Investigative Reporting. A Common Crawl-derived dataset alone included over two million documents from nytimes.com. The filings also describe employees actively working to bypass paywalls without detection, and deliberate efforts to strip copyright notices from training data before it reached the model.
This matters beyond just one lawsuit. Publishers including The Intercept, Raw Story, and others have filed similar claims. Licensing deals between AI companies and media organizations — like the ones OpenAI has struck with the Associated Press and Axel Springer — have been framed as voluntary partnerships. But if courts treat this evidence as proof that scraping caused real market harm, those deals start to look less like goodwill and more like legal risk management.
For developers building on top of OpenAI or Microsoft APIs, the bigger question is what a ruling against fair use would mean for training data practices across the industry. Anthropic, Google, and Meta all face versions of the same exposure. This case is the one most likely to produce precedent, and the internal admissions now on record make the companies’ public legal arguments harder to sustain.



