SpaceXAI's Grok 4.6 Scores 61 on Artificial Analysis Intelligence Index
Key point
Grok 4.6 achieved frontier-level intelligence and agent performance at a low cost.
Details
In the Artificial Analysis evaluation, Grok 4.6 scored 61 on the Intelligence Index, joining the frontier model group. This is the same score as GPT-5.6 Sol, trailing behind Claude Opus 5 (63 points) and Claude Fable 5 (62 points). It represents a 5-point increase over Grok 4.5 and a 23-point increase over Grok 4.3.
It showed particularly strong performance in agent tasks.
- GDPval-AA v2: Elo 1753, ranking second behind Claude Opus 5
- τ³-Banking: 50.7%, placing it in the top tier
- Terminal-Bench v2.1: 88.4%, comparable to leading models
- AA-Briefcase: Elo 1577, performing at the level of Claude Fable 5
API pricing remains the same as Grok 4.5: $2 per 1 million input tokens and $6 per 1 million output tokens. This is cheaper than Claude Opus 5 ($5/$25) and GPT-5.6 Sol ($5/$30), with a measured cost per task of $0.84. The price for cache-hit tokens is $0.5 per 1 million tokens.
The context window is 500K tokens. In AA-Briefcase, tasks were completed with an average of about 53 turns and 500 million input tokens, demonstrating higher token efficiency than Claude Opus 5 (max), which used about 103 turns and 2 billion input tokens.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.