Home
Several US Media Outlets Block Internet Archive's Wayback Machine Crawler to Prevent AI Training Abuse
As reported by Wired, several major media outlets and platforms—including The New York Times, Reddit, and USA Today's parent company—have recently blocked the Internet Archive's Wayback Machine. The move is intended to prevent AI companies from using the archive tool to indirectly scrape copyrighted material for model training.

The irony of blocking while benefiting
Ironically, USA Today's recent in-depth investigation into immigration policy statistics relied on historical data preserved by the Wayback Machine. Yet the media group's spokesperson stated they have fully blocked all crawlers, including ia_archiverbot, to address the growing risk of AI infringement.
How media organizations are applying restrictions
At least 23 mainstream news websites have now implemented restrictions:
Complete blocking: The New York Times and Reddit have directly blocked the Wayback Machine's dedicated crawler.
Interface filtering: The Guardian has not fully blocked crawlers but has excluded its content from the Internet Archive's API and filtered its search interface, making it extremely difficult for users to access historical archives.
In response to these blocking actions, over 100 active journalists, including Rachel Maddow, have co-signed a letter of support to the Electronic Frontier Foundation (EFF). They describe the Wayback Machine as an "indispensable tool" for fact-checking, tracking changes in the behavior of powerful institutions, and preserving digital history.
Publishers argue that AI companies training on the Internet Archive's vast data violates copyright law and directly competes with them. However, Internet Archive director Mark Graham warns that the ongoing closure of public web content is seriously undermining society's ability to understand historical truths and conduct public oversight. If this trend continues, a large portion of early digital records may be at risk of complete loss.
Related article
Claude Opus 5.2 Night Gray Launches with Faster Response, Solving Laziness Issue
Opus 5.2 quietly launched this morning, prompting many developers to notice that the updated model, Claude Opus 5.2, has begun a limited rollout within Claude Code.Yesterday evening, X users observed that invoking Opus 5 in Claude Code yielded perfor
Fireworks AI Unveils FireRouter with Opus: Cuts Encoding Costs by 57% with Minimal Accuracy Drop
Fireworks AI has introduced FireRouter with Opus, the industry’s first cache-aware routing system tailored for the Claude Opus series. Now accessible via a serverless endpoint, this independent routing model underwent over a month of internal A/B tes
NVIDIA Unveils Nemotron-Labs-Audex-30B-A3B Unified Audio Intelligence Model
As multimodal large models evolve rapidly, audio processing capabilities are frequently compromised—many models improve audio understanding at the expense of text logic. To address this, NVIDIA researchers have introduced Nemotron-Labs-Audex-30B-A3B
Related Special Topic Recommendations
Comments (0)
0/500
As reported by Wired, several major media outlets and platforms—including The New York Times, Reddit, and USA Today's parent company—have recently blocked the Internet Archive's Wayback Machine. The move is intended to prevent AI companies from using the archive tool to indirectly scrape copyrighted material for model training.

The irony of blocking while benefiting
Ironically, USA Today's recent in-depth investigation into immigration policy statistics relied on historical data preserved by the Wayback Machine. Yet the media group's spokesperson stated they have fully blocked all crawlers, including ia_archiverbot, to address the growing risk of AI infringement.
How media organizations are applying restrictions
At least 23 mainstream news websites have now implemented restrictions:
Complete blocking: The New York Times and Reddit have directly blocked the Wayback Machine's dedicated crawler.
Interface filtering: The Guardian has not fully blocked crawlers but has excluded its content from the Internet Archive's API and filtered its search interface, making it extremely difficult for users to access historical archives.
In response to these blocking actions, over 100 active journalists, including Rachel Maddow, have co-signed a letter of support to the Electronic Frontier Foundation (EFF). They describe the Wayback Machine as an "indispensable tool" for fact-checking, tracking changes in the behavior of powerful institutions, and preserving digital history.
Publishers argue that AI companies training on the Internet Archive's vast data violates copyright law and directly competes with them. However, Internet Archive director Mark Graham warns that the ongoing closure of public web content is seriously undermining society's ability to understand historical truths and conduct public oversight. If this trend continues, a large portion of early digital records may be at risk of complete loss.
Claude Opus 5.2 Night Gray Launches with Faster Response, Solving Laziness Issue
Opus 5.2 quietly launched this morning, prompting many developers to notice that the updated model, Claude Opus 5.2, has begun a limited rollout within Claude Code.Yesterday evening, X users observed that invoking Opus 5 in Claude Code yielded perfor
Fireworks AI Unveils FireRouter with Opus: Cuts Encoding Costs by 57% with Minimal Accuracy Drop
Fireworks AI has introduced FireRouter with Opus, the industry’s first cache-aware routing system tailored for the Claude Opus series. Now accessible via a serverless endpoint, this independent routing model underwent over a month of internal A/B tes
NVIDIA Unveils Nemotron-Labs-Audex-30B-A3B Unified Audio Intelligence Model
As multimodal large models evolve rapidly, audio processing capabilities are frequently compromised—many models improve audio understanding at the expense of text logic. To address this, NVIDIA researchers have introduced Nemotron-Labs-Audex-30B-A3B











