Artificial Analysis Releases Capability Indices v1.1… Claude Fable 5.1 Tops All 6 Categories, Grok 4.5 Noted for Cost Efficiency
Key point
Artificial Analysis released Capability Indices v1.1, with Claude Fable 5.1 taking first place in all six occupational indices and Grok 4.5 demonstrating high cost efficiency.
Details
Artificial Analysis released Capability Indices v1.1 and Intelligence Index v4.2, re-evaluating the performance and cost of the latest AI models. Claude Fable 5.1 ranked first in all six Capability Indices: Finance & Accounting, Strategy & Ops, Legal, Healthcare & Medical, Engineering, and Economics. Among open-weight models, Kimi K3 took first place in Finance & Accounting, Legal, and Economics; DeepSeek V4.1 Flash in Strategy & Ops; and GLM-5.3 in Healthcare & Medical and Engineering.
Grok 4.5 demonstrated strong performance in agentic tasks, ranking 4th (Elo 1543) on GDPval-AA v2 and scoring 33% on 𝜏³-Banking, surpassing GPT-5.5 (31%). In terms of cost, it recorded a headline price of $0.31 per task based on the Intelligence Index and $2.49 based on the Coding Agent Index (Grok Build), making it over 60% cheaper than Opus 4.8 ($11.80) and positioning it on the cost-performance Pareto frontier.
Gemini 3.1 Pro Preview ranked first in 6 of the 10 evaluations comprising the Intelligence Index (including Terminal-Bench Hard, AA-Omniscience, etc.) and can run at less than half the cost of leading models such as Opus 4.6 and GPT-5.5. However, it costs approximately twice as much as leading open-weight models like GLM-5. Additionally, the newly released Endpoint Accuracy Index measures the accuracy preservation rates of third-party hosting for GLM-5.2, gpt-oss-120b, and DeepSeek V4 Pro.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.