Why Measuring AI Performance Is Getting Harder
Key point
As AI model performance improves rapidly, measurement errors in existing benchmarks are growing and gaps with real-world environments are emerging, making performance measurement increasingly difficult.
Details
The METR chart, a representative indicator showing the pace of AI model advancement, compares model performance by converting the complexity of software engineering tasks into human working time. Starting from GPT-3.5, which could perform tasks worth 30 seconds, it is estimated that the recent Claude Opus 4.6 can perform tasks that take humans 12 hours.
However, this rapid performance improvement may be a statistical illusion. Claude Opus 4.6 has been solving the hardest problems within the METR test set, making it difficult to determine an upper bound on performance, and as a result the confidence interval has widened significantly, ranging from 5 hours to 66 hours, increasing measurement uncertainty.
A more fundamental problem is that existing benchmarks only measure well-defined, independent tasks. Real-world work is connected to other tasks, requires collaboration with others, and has the complex characteristic of goals that shift fluidly.
As AI comes to perform long-term tasks spanning weeks or months, beyond tasks measured in hours, current measurement methods will have limits in evaluating actual capability.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.