PaperBench: A Benchmark for Evaluating AI Research Reproduction Ability
Key point
OpenAI has released PaperBench, a benchmark that evaluates AI agents' ability to reproduce state-of-the-art AI research.
Details
PaperBench is a benchmark that evaluates how well AI agents can reproduce state-of-the-art AI research. Agents must, for 20 Spotlight and Oral papers from ICML 2024, understand the papers' contributions, develop a codebase, and successfully run experiments from start to finish.
To ensure evaluation accuracy, rubrics developed in collaboration with each paper's authors are used, comprising a total of 8,316 detailed evaluation items. In addition, an LLM-based judge model was developed to automatically grade reproduction attempts based on the rubrics, enabling large-scale evaluation.
Key evaluation results are as follows:
- The best-performing model, Claude 3.5 Sonnet (New) (when using an open-source scaffold), achieved an average replication score of 21.0%.
- Testing against top ML PhD-level talent showed that current AI models have not yet reached human-level performance.
OpenAI has open-sourced the related code to support future research aimed at understanding the engineering capabilities of AI agents.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.