AI Briefing
KO

News Outlets Block Wayback Machine

·2026.05.01 20:32

Key point

23 major news publishers have blocked Wayback Machine crawlers over AI training concerns.

Details

Major news outlets are blocking the Internet Archive's Wayback Machine crawler.

According to Originality AI analysis, 23 major news publishers are blocking ia_archiverbot, and 241 news sites across 9 countries have restricted access to at least one Archive bot. This includes The New York Times, CNN, USA Today, and The Guardian, with The New York Times applying a full block since late 2025. The Guardian avoided a full block by coordinating directly with the Internet Archive to adjust access restrictions.

News outlets believe AI companies are using archived articles to train large-scale models at scale, competing with the original content. The Internet Archive counters that it already operates bulk-download limits and controls on automated extraction, and that the issue isn't the Archive's collection itself but how third parties use the archived data.

The side effects of the blocking are significant.

  • The Wayback Machine has been used to verify article edit histories, submit evidence in court, and preserve Wikipedia links.
  • Wikipedia has linked to 2.6 million+ news articles preserved by the Wayback Machine.
  • EFF argues that AI companies should be targeted directly, and public records shouldn't be sacrificed in the process.
  • The Fight for the Future petition has been signed by more than 100 working journalists.

Ultimately, the dispute over controlling AI training data is now putting pressure on public web archives as well.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.