Evaluating Scientific Discovery Agents (4 min read)
Key point
ScienceWorld and DiscoveryWorld reveal how AI science agents actually perform.
Details
AI science agents are appearing everywhere, but whether they can truly do science is a separate matter. Ai2 built benchmarks called ScienceWorld and DiscoveryWorld to verify this, and their significance is growing again as recent models reach these tasks.
ScienceWorld is an environment that tests whether models can actually perform elementary-school-level science experiments. In 2022, even top-tier models scored high on multiple-choice science tests, but when asked to translate that same knowledge into performing experiments, they scored below 10%. In contrast, as of early 2025, top frontier models have climbed to the low 80s, though they still haven't fully mastered the 4th-grade science curriculum.
DiscoveryWorld is a step harder. On the fictional space colony Planet X, it has agents solve 120 tasks such as identifying the cause of an infection and inferring mathematical relationships in a quantum reactor, spanning 8 topics, 3 difficulty levels, and parametric variations that differ each run. Agents must form hypotheses, design experiments, execute them, and analyze results, and they're evaluated not just on whether they get the right answer but on whether they followed scientific procedure.
The key comparison is as follows.
- ScienceWorld: measures whether textbook-level science concepts can be reproduced through experiments
- DiscoveryWorld: measures whether new discoveries can be designed and verified from scratch
- Experimental results: even the best recent systems solve only about 20% of DiscoveryWorld's hardest tasks
- Human baseline: an average highly educated scientist solves about 70%
These two benchmarks reveal the gap between "knowing knowledge" and "actually performing scientific discovery." Ai2 emphasizes that for science agents to lead to disease treatments, new materials, and new discoveries, they must first prove their actual performance on these open-ended, end-to-end scientific discovery tasks rather than closed-form ones.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.