Audit of 162 AI Benchmark Gaps Finds Only 20 Separate Cleanly
Key point
An independent audit of launch charts and leaderboards reveals that most claimed performance wins between Claude, GPT, Gemini, and Mistral models do not separate cleanly from statistical noise.
Details
An analysis of 162 benchmark gaps from recent launches of Claude, GPT, Gemini, and Mistral models, along with 9 leaderboards, found that only 20 gaps separated cleanly when accounting for sampling noise. The study highlights a critical issue in how AI model performance is reported and compared, suggesting that many widely cited "wins" are statistically indistinguishable or unverifiable.
Launch Chart Reliability
The audit examined 44 gaps quoted in six frontier launch posts. Of these:
- 11 separated cleanly.
- 5 did not separate.
- 28 were not checkable from the published data due to issues like averaged runs with unstated repetition counts, mismatched harnesses, or unclear effort settings.
Specifically, Gemini 4 Argon had 15 quoted comparisons where public numbers did not support an honest interval, and Mistral Large 4 had 9. None of these 24 comparisons could be honestly verified from the published data.
Leaderboard Noise
The analysis extended to 118 adjacent pairs across nine leaderboard views, of which 9 separated. A notable example is SWE-bench Multilingual, where the top two scores (72.7% vs 66.3%) appear decisive. However, with only 300 tasks, the statistical resolution is low. The confidence interval for this specific pair includes zero, meaning the difference is not statistically significant. In fact, six positions on that leaderboard cannot be distinguished from the first place.
Methodology and Transparency
The author initially identified 21 separated gaps but revised the count to 20 after a rigorous audit against benchmark score types and published uncertainty metrics. The analysis pipeline, built with Claude Code, is public and reproducible. The data, write-up, and an interactive calculator for testing benchmark gaps are available at driftproofhq.com.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.