On September 17, The New York Times and other news organizations filed unredacted motions for summary judgment in their copyright case against OpenAI and Microsoft, exposing internal comments the companies had fought to keep confidential. The suit, filed in December 2023, has been consolidated into multi-district litigation before a single federal judge and is now at the summary-judgment stage. The question at its center: whether training models on copyrighted news articles is fair use.
[1]The sharpest language comes from inside Microsoft. Brent Hecht, the company's director of applied science, wrote in a January 2023 memo that copying millions of news articles into large language models without permission could amount to "an astonishing theft of unprecedented proportions" and "the largest theft of labor in human history." A January 2024 internal deck by Hecht warned of a "doom loop": Microsoft's own data showed that users of its AI answer tools clicked through to publisher pages 83% to 93% less than users of traditional search. A Microsoft document quoted in the filing puts the relationship bluntly: "It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its 'content supply chain.'" Microsoft responded Thursday that the comments reflect "one employee's individual perspective" and are "not a legal analysis."
[1][2]OpenAI's side is pinned down by its own paper trail. Nick Turley, head of ChatGPT, wrote that AI products pose an "existential threat" to publishers and are "largely substitutive," adding "that's just what it is, and they're going to get more and more substitutive as they get better." President Greg Brockman wrote around 2017 that he was "deeply driven by astronomical wealth" from commercializing OpenAI's technology. When researcher Nick Ryder told him about a "hack to get around nytimes paywall," Brockman replied: "ah nice."
The scale figures are now public for the first time. OpenAI's mid-training datasets contain more than 91,692 copies of works by the Times, the Daily News and the Center for Investigative Reporting; a Common Crawl-derived dataset holds more than 2 million documents from nytimes.com alone; and Project Mango, a data-sharing initiative with Microsoft, includes at least 160,903 unique works from news publishers. Microsoft CEO Satya Nadella testified that paywalled content should be licensed for AI training and that, had he known OpenAI scraped paywalled material, he would have invoked Microsoft's contractual right to demand retraining.
[1][2]Fair use's fourth factor — whether a use substitutes for the original and harms its market — is exactly what these documents speak to. Plaintiff lawyer Steven Lieberman said the disclosures show OpenAI and Microsoft "knowingly committed theft" rather than operating within fair use. The Trump administration filed an amicus brief on September 1 arguing AI training is "highly transformative"; the same day, roughly 400 local U.S. newspapers filed a class action against the two companies.
None of this decides the case — fair-use disputes survive on discovery, experts and appeals. But the Times now holds what plaintiffs rarely get: the defendants' own executives, in writing, describing the practice at the center of the suit as theft, and their own products as substitutes for the journalism they were built on. In a courtroom weighing market harm, that paper trail may matter more than any outside witness.
[1][2]