NYT Lawsuit: What Internal Statements at Microsoft and OpenAI Reveal

Zeitungen und ein Richterhammer als Symbol für den Urheberrechtsstreit um KI-Training
Photo by Sasun Bughdaryan on Unsplash

On September 17, 2026, a previously redacted filing was unsealed in The New York Times’ copyright case against OpenAI and Microsoft. It quotes executives at both companies on what they wrote internally about training AI models – and the language sounds far less relaxed than their public defense. On the key legal question, whether such training counts as fair use, the statements could carry weight.

Key takeaways

  • The unsealed brief, with which the Times seeks summary judgment (a ruling without a full trial), quotes internal statements and sworn testimony from Microsoft and OpenAI executives.
  • Microsoft’s director of applied science, Brent Hecht, called mass scraping of content “an astonishing theft of unprecedented proportions” as early as January 2023; OpenAI’s head of ChatGPT, Nick Turley, wrote that the products are “largely substitutive.”
  • According to the Times’ brief, Copilot cut click-through rates to the newspaper’s domain by as much as 93 percent compared with classic Bing search.
  • The quotes come from the plaintiff’s brief; the underlying exhibits remain sealed, and OpenAI and Microsoft have not commented so far.
  • There is no ruling. But the statements show where the Times sees its strongest lever: whether AI products replace the market for the originals.

What the brief says

The Times sued OpenAI and Microsoft at the end of 2023, alleging that they used millions of its articles without permission to train their language models. The case is heard in federal court in New York and is one of the central precedents on whether AI training falls under fair use. That is the U.S. doctrine allowing the use of copyrighted works without a license under certain conditions. According to 404 Media, the unredacted brief became visible on the CourtListener platform on September 17; passages that had been redacted at the request of Microsoft and OpenAI are now readable.

Hecht’s wording is the starkest. In an internal memo from January 2023, the Microsoft director wrote, according to TechCrunch, that scraping content for AI training was “the largest theft of labor in human history.” The brief presents this as a warning about how “millions of people” might soon see such practices – not as a legal assessment of whether Microsoft infringed copyright. That distinction is worth keeping in mind.

In an internal presentation from January 2024, Hecht also described a “doom loop”: if AI answers cut off traffic to the sources, that hurts the performance of the models and the entire web at the same time. It is “highly unusual,” he wrote, for an end product to threaten the economic foundation of its own suppliers. Another Microsoft document, whose author is not named, warns that generative AI could “significantly disrupt” the employment of the very people who created the training data.

Substitution instead of supplement

Legally more delicate than any quote about “theft” is a different finding: according to the brief, both companies concede that their products replace originals. Nick Turley, who runs ChatGPT at OpenAI, wrote internally that publishers faced an “existential threat,” that the products were “largely substitutive,” and that they would become more so with every model generation. heise online also reports that Microsoft CEO Satya Nadella confirmed under oath that chatbot conversations replace a visit to the original source.

Then there is a number: according to the Times’ brief, Copilot’s answer engine cut click-through rates to the newspaper’s domain by as much as 93 percent compared with classic Bing search. 404 Media cites testimony from Microsoft managers putting the drop at “more than 90 percent.” An OpenAI engineer reportedly testified that users do not click, no matter how prominently links are shown.

The structure of the fair-use test explains why this matters. Courts weigh, among other things, whether a use is transformative and how it affects the market for the original. If a company itself writes that its product replaces the source, it hands the other side material for exactly that point. TechCrunch notes that courts have so far been largely receptive to the AI companies’ fair-use arguments; the new passages weaken their case on substitution and market harm.

Access to paid content and datasets

Besides the executive quotes, the brief contains allegations about the data itself. OpenAI’s training data is said to include more than 91,000 copies of works from the New York Times, the Daily News and the Center for Investigative Reporting, and a dataset derived from Common Crawl allegedly held more than two million documents from nytimes.com. Microsoft is alleged to have passed data to OpenAI through projects called “Taxi” and “Mango”; the Mango dataset reportedly contains at least 160,903 individual works from the publishers. These are the plaintiff’s allegations, not findings by a court.

The paywall also plays a role. When researcher Nick Ryder told OpenAI president Greg Brockman about a “hack” to get around the Times’ paywall, he reportedly replied “ah nice.” Nadella, in turn, said in his deposition, according to TechCrunch, that anything paywalled should be licensed by anyone who wants to use it. Had he known that OpenAI trained on paywalled content, he would have used his right to demand that the models be retrained.

Assessment: what this proves – and what it does not

Three caveats belong to an honest reading. First, the quotes come from the brief of a party that assembles its evidence for maximum effect; the original documents remain closed to the public, and quotes stripped of context can mislead. Second, an internal warning does not mean a court will treat the conduct as a violation of the law. Third, neither OpenAI nor Microsoft has answered requests for comment, so their side is missing here.

Even so, the brief shifts the debate. Until now the defense rested mainly on the claim that training is a transformative process. Now it faces the executives’ own statements that the products replace publishers’ offerings – and that Microsoft warned internally of a loop that dries up its own data supply. How unusual the political situation is becomes clear from the other side: on September 2, the U.S. Justice Department had sided with OpenAI in a “statement of interest.”

Outlook

With the motion for summary judgment, the question now lies with Judge Sidney Stein, who must decide whether the facts are clear enough that no trial before a jury is needed. The reports name no schedule, and the outcome is open. For readers and publishers, one sober insight emerges for now: the companies evidently know very well that AI answers cost clicks. Whether a court treats that knowledge as an admission of guilt, a market observation or a footnote will decide whether publishers can demand licenses in the future or simply keep losing traffic.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top