AI Briefing
KO

AI evals become the new compute bottleneck

·2026.05.01 18:52

Key point

AI eval costs are surging, making evaluation a new compute bottleneck.

Details

Evaluation costs grow sharply as you move from static benchmarks to agentic benchmarks to training-in-the-loop.

  • HELM cost $85~$10,926 in API cost per model, open models required 540~4,200 GPU-hours, and the combined cost across all 30 models and 42 scenarios was roughly $100,000.

  • tinyBenchmarks shrank MMLU from 14,000 items to 100 while preserving rankings within about 2% error, and the Open LLM Leaderboard was also reduced from 29,000 to 180 items. Static benchmarks still allowed large-scale subsampling.

  • HAL spent about $40,000 on 21,730 rollouts, and as of April 2026 total rollouts had grown to 26,597. The cost of a single run varied from $0.12~$2,829 depending on the benchmark, and model × scaffold × token budget combinations swung cost by more than 10x.

  • In Mind2Web and GAIA, more expensive setups did not guarantee better accuracy, and CLEAR found that the accuracy-optimal configuration was 4.4~10.8x more expensive than Pareto-efficient alternatives.

  • The Well costs about 960 H100-hours (roughly $2,400) per new architecture, and the full sweep is 3,840 H100-hours (roughly $9,600).

  • When evaluation involves actual training, as with MLE-Bench, RE-Bench, ResearchGym, and PaperBench, a single evaluation turns into tens to hundreds of GPU-hours or thousands of dollars. PaperBench full costs about $9,500, Code-Dev about $4,200, and NAS-Bench-101 required over 100 TPU-years for tabular construction.

Running repeated trials for reliability multiplies costs even further. Ultimately, as benchmarks move closer to real tasks, the compression that was possible at 100~200x for static benchmarks drops to 2~3.5x for agentic benchmarks and to nearly 1x for training-in-the-loop.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.