AI Scientists Produce Results Without Scientific Reasoning
·2026.04.23 07:14
Key point
Across 25,000 experiments, AI scientists ignored evidence and failed to update their beliefs.
Details
An evaluation of LLM-based AI scientists across 8 domains and more than 25,000 runs found that the self-correcting pattern of scientific inquiry rarely emerged.
- Even after collecting evidence, 68% ignored it.
- Hypotheses were changed in only 26% of cases when disconfirming data appeared.
- Behavioral patterns did not differ significantly between workflow execution and hypothesis-driven inquiry.
In the performance breakdown, the base model accounted for 41.4% of explained variance, while scaffold influence was only 1.5%. The conclusion is that scaffolding improvements alone—such as ReAct, structured tool-calling, and chain-of-thought—cannot solve this problem.