METR AI Progress Graph Criticized for Errors
Key point
The METR graph, a key piece of evidence in discussions of AI's rapid progress, has come under criticism for containing serious methodological errors.
Details
The METR Long Tasks benchmark graph, frequently cited as evidence supporting AI's rapid advancement, has been called unreliable.
Nathan Witkin, a research writer at NYU Stern, criticized the graph for containing compounding errors that make it impossible to draw meaningful conclusions. The main issues are as follows.
- Sample generalization error: Data collected from a small number of the researchers' colleagues was generalized to represent the full dataset.
- Unverified human baselines: Some human baseline data was not actually measured but was instead estimated by the authors.
- Data bias: By paying human benchmarkers hourly wages, the data on task completion time was distorted by the incentive structure.
Witkin argued that while the graph looks overly smooth and sophisticated, it contains so many errors that it should be discarded in favor of seeking higher-quality information.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.