Home
Microsoft exec calls AI data scraping 'biggest theft of labor in history,' unredacted filings show

Unsealed documents from the three-year-old copyright lawsuit filed by The New York Times against OpenAI and Microsoft reveal an admission that AI scraping constitutes theft and poses a severe threat to media outlets.
According to the lawsuit, a senior Microsoft executive privately labeled the companies’ AI training methods as “theft,” while OpenAI’s leadership acknowledged that its models created an “existential threat” to the publishers and journalists whose content trained them.
The newly unsealed records also detail how the companies allegedly accessed and utilized this content by bypassing paywalls undetected, constructing training datasets through mass scraping, and intentionally removing copyright notices from the data.
It is important to note that much of this new information originates from The Times’ own legal brief rather than the underlying exhibits, which remain sealed. The quotes below are presented without their original context.
This unredacted filing marks the latest escalation in the three-year-old lawsuit, where The New York Times initially claimed the firms violated copyright law by training generative AI models on its content.
There is no clear legal answer regarding whether AI firms can legally use copyrighted material for training, though judges have generally favored the AI companies’ argument that such training qualifies as “fair use.” This legal doctrine permits the use of copyrighted works without permission in specific instances, such as parody, news reporting, or criticism. Earlier this month, the Trump administration submitted a brief supporting OpenAI’s unlicensed use of copyrighted material to train its large language models.
However, several of these new admissions contradict OpenAI’s fair use defense, particularly the requirement that such use does not substitute for or harm the market for the original work.
For instance, Microsoft’s own data indicates that its Copilot “answer engine” caused click-through rates for The New York Times’ domain to plummet by up to 93% compared to traditional Bing searches. An internal Microsoft presentation authored by Brent Hecht, Microsoft’s Director of Applied Science, in January 2024, describes this decline as a “doom loop” that would simultaneously “hurt the performance of our models and the entire web.”
“It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its ‘content supply chain,’” the Microsoft document states, as quoted in the filing.
Microsoft CEO Satya Nadella also testified in a deposition earlier this year that “anything that is paywalled should be licensed by anyone who wants to use it…for grounding or training,” and clarified that had he “been made aware that OpenAI had scraped and trained on information that was behind a paywall,” he would have “invoked [Microsoft’s right to] require OpenAI to retrain its models.”
Other admissions undermine different aspects of the fair-use test: OpenAI’s Head of ChatGPT, Nick Turley, wrote in internal communications that publishers face an “existential threat” from products like the chatbot, which are “largely substitutive” and “will get more and more substitutive as they get better.”
OpenAI President Greg Brockman described the models as “excellent at news.” Nadella agreed under oath earlier this year that interacting with chatbots “has substituted … giving you the information right there on the website on the AI platform versus needing to go to the underlying source.”
Such language highlights how the technology could directly compete with, rather than transform, the original work.
A Microsoft document states there is a “real risk” that generative AI could “significantly disrupt the employment of the very people who generated the data on which the foundation model was trained.”
The sheer scale of the copying is striking. The documents reveal for the first time that OpenAI’s mid-training datasets alone contain more than 91,692 copies of works published by the NYT, Daily News, and Center for Investigative Reporting. A dataset derived from Common Crawl included more than 2 million documents from nytimes.com alone.
In a January 2023 internal memo, Hecht called it “an astonishing theft of unprecedented proportions” and “the largest theft of labor in human history.”
The filing provides new details on how OpenAI and Microsoft acquired the plaintiffs’ content, including scraping it from the Bing Index.
“OpenAI delivered the entire GPT-3 training dataset to Microsoft, which Microsoft used to evaluate how to implement OpenAI’s models within its own commercial products,” the filing states. “Microsoft similarly provided training data to OpenAI through initiatives called Project Taxi and Project Mango.”
The companies allegedly assembled the Project Mango data into a training dataset containing copies of at least 160,903 unique works from news publishers.
To maximize the effectiveness of their scraping, OpenAI employees allegedly devised a plan to circumvent paywalls without detection. The filings show that when OpenAI researcher Nick Ryder informed Brockman about a “hack to get around nytimes paywall,” Brockman replied: “ah nice.”
OpenAI employees also allegedly constructed training datasets like WebText and WebText2 that disproportionately relied on scraped news content. They also allegedly extracted millions of articles from Common Crawl, a free, open repository of web crawl data. The findings also describe deliberate efforts to strip copyright notices from training data before it reached the model, since researchers “wouldn’t want model outputting” “copyright notices” to users.
OpenAI and Microsoft did not return requests for comment.
Related article
OpenAI GPT-6 Halves Token Costs Across Benchmarks as Sol and Luna Join
According to Artificial Analysis, GPT-6 Sol (max) ranks as the top model for intelligence, priced at $48. Image credit: Getty ImagesAfter OpenAI released GPT-6 Sol and Luna, Matt Weaver, Head of Solutions Engineering, explained to AI Magazine how the
The AI graveyard: A running list of projects and startups that didn’t make it
Relay, an AI-driven workflow automation platform positioned as a Zapier alternative, ceased operations on Monday. The service enabled users to automate email and task workflows via AI agents, but as OpenAI, Google, and other major platforms integrate
OpenAI bets on families as ChatGPT goes deeper into households
Over three years since ChatGPT propelled generative AI into the spotlight, OpenAI is expanding its scope from individual users to entire households.OpenAI is recruiting a product manager in San Francisco to develop family-oriented experiences across
Related Special Topic Recommendations
Comments (0)
0/500

Unsealed documents from the three-year-old copyright lawsuit filed by The New York Times against OpenAI and Microsoft reveal an admission that AI scraping constitutes theft and poses a severe threat to media outlets.
According to the lawsuit, a senior Microsoft executive privately labeled the companies’ AI training methods as “theft,” while OpenAI’s leadership acknowledged that its models created an “existential threat” to the publishers and journalists whose content trained them.
The newly unsealed records also detail how the companies allegedly accessed and utilized this content by bypassing paywalls undetected, constructing training datasets through mass scraping, and intentionally removing copyright notices from the data.
It is important to note that much of this new information originates from The Times’ own legal brief rather than the underlying exhibits, which remain sealed. The quotes below are presented without their original context.
This unredacted filing marks the latest escalation in the three-year-old lawsuit, where The New York Times initially claimed the firms violated copyright law by training generative AI models on its content.
There is no clear legal answer regarding whether AI firms can legally use copyrighted material for training, though judges have generally favored the AI companies’ argument that such training qualifies as “fair use.” This legal doctrine permits the use of copyrighted works without permission in specific instances, such as parody, news reporting, or criticism. Earlier this month, the Trump administration submitted a brief supporting OpenAI’s unlicensed use of copyrighted material to train its large language models.
However, several of these new admissions contradict OpenAI’s fair use defense, particularly the requirement that such use does not substitute for or harm the market for the original work.
For instance, Microsoft’s own data indicates that its Copilot “answer engine” caused click-through rates for The New York Times’ domain to plummet by up to 93% compared to traditional Bing searches. An internal Microsoft presentation authored by Brent Hecht, Microsoft’s Director of Applied Science, in January 2024, describes this decline as a “doom loop” that would simultaneously “hurt the performance of our models and the entire web.”
“It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its ‘content supply chain,’” the Microsoft document states, as quoted in the filing.
Microsoft CEO Satya Nadella also testified in a deposition earlier this year that “anything that is paywalled should be licensed by anyone who wants to use it…for grounding or training,” and clarified that had he “been made aware that OpenAI had scraped and trained on information that was behind a paywall,” he would have “invoked [Microsoft’s right to] require OpenAI to retrain its models.”
Other admissions undermine different aspects of the fair-use test: OpenAI’s Head of ChatGPT, Nick Turley, wrote in internal communications that publishers face an “existential threat” from products like the chatbot, which are “largely substitutive” and “will get more and more substitutive as they get better.”
OpenAI President Greg Brockman described the models as “excellent at news.” Nadella agreed under oath earlier this year that interacting with chatbots “has substituted … giving you the information right there on the website on the AI platform versus needing to go to the underlying source.”
Such language highlights how the technology could directly compete with, rather than transform, the original work.
A Microsoft document states there is a “real risk” that generative AI could “significantly disrupt the employment of the very people who generated the data on which the foundation model was trained.”
The sheer scale of the copying is striking. The documents reveal for the first time that OpenAI’s mid-training datasets alone contain more than 91,692 copies of works published by the NYT, Daily News, and Center for Investigative Reporting. A dataset derived from Common Crawl included more than 2 million documents from nytimes.com alone.
In a January 2023 internal memo, Hecht called it “an astonishing theft of unprecedented proportions” and “the largest theft of labor in human history.”
The filing provides new details on how OpenAI and Microsoft acquired the plaintiffs’ content, including scraping it from the Bing Index.
“OpenAI delivered the entire GPT-3 training dataset to Microsoft, which Microsoft used to evaluate how to implement OpenAI’s models within its own commercial products,” the filing states. “Microsoft similarly provided training data to OpenAI through initiatives called Project Taxi and Project Mango.”
The companies allegedly assembled the Project Mango data into a training dataset containing copies of at least 160,903 unique works from news publishers.
To maximize the effectiveness of their scraping, OpenAI employees allegedly devised a plan to circumvent paywalls without detection. The filings show that when OpenAI researcher Nick Ryder informed Brockman about a “hack to get around nytimes paywall,” Brockman replied: “ah nice.”
OpenAI employees also allegedly constructed training datasets like WebText and WebText2 that disproportionately relied on scraped news content. They also allegedly extracted millions of articles from Common Crawl, a free, open repository of web crawl data. The findings also describe deliberate efforts to strip copyright notices from training data before it reached the model, since researchers “wouldn’t want model outputting” “copyright notices” to users.
OpenAI and Microsoft did not return requests for comment.
OpenAI GPT-6 Halves Token Costs Across Benchmarks as Sol and Luna Join
According to Artificial Analysis, GPT-6 Sol (max) ranks as the top model for intelligence, priced at $48. Image credit: Getty ImagesAfter OpenAI released GPT-6 Sol and Luna, Matt Weaver, Head of Solutions Engineering, explained to AI Magazine how the
The AI graveyard: A running list of projects and startups that didn’t make it
Relay, an AI-driven workflow automation platform positioned as a Zapier alternative, ceased operations on Monday. The service enabled users to automate email and task workflows via AI agents, but as OpenAI, Google, and other major platforms integrate
OpenAI bets on families as ChatGPT goes deeper into households
Over three years since ChatGPT propelled generative AI into the spotlight, OpenAI is expanding its scope from individual users to entire households.OpenAI is recruiting a product manager in San Francisco to develop family-oriented experiences across











