AI Briefing
KO

Stanford Researchers Release Terminal-Bench-Science 0.1, an AI Agent Benchmark Based on Scientific Workflows

·2026.08.28 09:00

Key point

Stanford researchers have released Terminal-Bench-Science 0.1, an AI agent benchmark based on real scientific workflows.

Details

The AI research team at Stanford University and the Terminal-Bench development team collaborated to release Terminal-Bench-Science 0.1, an AI agent evaluation benchmark based on real scientific research workflows. This benchmark establishes standards for measuring AI's scientific capabilities, led by actual scientists rather than model developers or data providers.

Features and Goals of the Benchmark

Terminal-Bench-Science is designed to evaluate the ability of AI agents to perform complex scientific tasks. Key features include:

  • Based on Real Workflows: Includes 70 tasks extracted from actual research settings, rather than textbook questions or standardized exercises.
  • Verifiable Results: Evaluates concrete outputs such as analysis, simulations, proofs, code, and data artifacts through reproducible tests.
  • Continuous Evolution: The benchmark evolves alongside advancements in AI technology, forming a feedback loop between the scientific community and AI development.

Evaluation Results and Task Scope

The initial release includes 70 tasks across life sciences, physics, earth sciences, mathematics, and engineering. Among the models evaluated so far, Claude Opus 5 achieved the highest performance, recording a 30% resolution rate on Terminal-Bench-Science 0.1. This suggests that AI agents still have limitations in fully autonomously solving complex scientific problems.

The benchmark aims to contribute to automating technically demanding and time-consuming workflows, allowing AI agents to focus on aspects requiring human judgment, such as defining research questions, formulating hypotheses, and interpreting results.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.