AI Briefing
KO

AI evals are becoming the new compute bottleneck

·2026.04.30 02:00

Key point

Static benchmarks keep getting compressed, but agent and training evals are seeing costs skyrocket.

1 / 2

Details

Static benchmarks can still be compressed by as much as 100–200x, but as you move toward agent benchmarks and training-in-the-loop evals, the cost savings shrink sharply.

  • HAL spent about $40,000 on 21,730 rollouts across 9 models and 9 benchmarks, with a single GAIA run reaching up to $2,829 for frontier models.
  • Exgentic's $22,000 sweep showed a 33x cost difference on the same task, revealing that scaffold choice is the primary cost driver.
  • Even within the same benchmark, costs can diverge by 3–4 orders of magnitude depending on the model × scaffold × token budget combination.
  • The Well requires 960 H100-hours (about $2,400) to evaluate a single new architecture, and 3,840 H100-hours (about $9,600) for a full 4-baseline sweep.
  • Benchmarks closer to real-world tasks, like PaperBench, MLE-Bench, ResearchGym, and RE-Bench, see evaluation costs soar into the thousands to tens of thousands of dollars.

Securing reliability through repeated runs drives costs even higher. As benchmarks move closer to real-world tasks, evaluation itself becomes a new compute bottleneck.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.