AI Briefing
KO

The Truth Behind AI Agent Performance Gains: Algorithms vs Models

·2026.06.01 23:34

Key point

An analysis suggesting that MLE-Bench performance gains stem more from improved model performance and increased search volume than from algorithmic advances has been released, alongside a new benchmark called FML-Bench.

Details

Over the past two years, MLE-Bench scores have surged from 30% to 80%, and an analysis has been presented examining whether this achievement reflects genuine algorithmic advancement or is simply the result of better base models, increased search volume, or overfitting.

When controlling for the same step budget and model, the AIDE algorithm from two years ago was found to perform on par with the latest agents and evolutionary search systems. This suggests that recent performance gains rely more heavily on model improvements and expanded search volume than on innovation in the algorithms themselves.

Alongside this, a new automated ML research benchmark called FML-Bench has been released to precisely measure agents' algorithmic efficiency (search and memory). FML-Bench is designed to evaluate the pure efficiency of algorithms by integrating a code-editing agent, step definitions, and verification/test splits.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.