AI Briefing
KO

Reproducing 2,200 ICML Papers Using AI Agents

·2026.08.13 09:00

Key point

A large-scale effort to reproduce over 2,200 ICML papers using AI agents verified the reliability of research.

Details

HuggingFace conducted a large-scale experiment to reproduce 2,226 ICML papers (approximately 34% of the entire conference) using AI agents through a hackathon involving more than 1,200 community members.

Experimental Method and Tools

  • Participants used coding agents such as Claude Code, Cursor, and Codex to extract key claims from the papers, write and execute experimental code, and report results.
  • All experimental processes were recorded in the Trackio logbook, and the open-weight model GLM-5.2 served as an automated reader to determine whether each claim was Verified, Falsified, Toy, or Inconclusive.

Key Results

  • 51% of papers: At least one claim was independently verified.
  • 23% of papers: Claims were found to be false or controversial.
  • Notably, in 242 papers, different reproduction teams reached conflicting verdicts on the same claims, demonstrating that reproducibility is not a simple binary issue.

This experiment suggests that, in response to the surge in AI research papers, AI agents can contribute to ensuring research reliability by automating the verification and reproduction processes.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.