Artificial Analysis Releases Intelligence Index v4.3
Key point
Artificial Analysis released Intelligence Index v4.3, introducing Terminal-Bench 4.0 and AutomationBench-AA.
Details
Artificial Analysis released Intelligence Index v4.3, which increases the difficulty of agentic coding and expands workflow types. This update serves as a transitional step before v5, focusing on increasing the proportion of private test sets to better reflect real-world problem-solving capabilities and prevent evaluation manipulation.
Key Changes
- Terminal-Bench 2.1 → 4.0 Upgrade: Includes 66 terminal-based multi-step tasks across software engineering, machine learning, science, and operations. Computational and time limits were recalibrated, and verification methods were improved. The harness was changed from Terminus 2 to the model-agnostic mini-SWE-agent.
- Introduction of AutomationBench-AA: Replaces the previous 𝜏³-Banking with Zapier's business workflow automation benchmark. It tests 657 business workflows in simulated applications such as Gmail, Slack, and Salesforce.
Evaluation Weight Adjustments
The proportion of evaluations using private test sets or gold answers increased from 40% to 45%. This measure aims to prevent benchmark score gaming. Category weights remain the same as in v4.2:
- Agents: 30%
- Coding: 20%
- General: 30%
- Scientific Reasoning: 20%
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.