AI Briefing
KO

Serious Flaws Pointed Out in METR's AI Time Horizon Graph

·2026.05.26 03:30

Key point

Criticism has emerged that METR's AI time horizon graph has serious flaws, including data bias and measurement errors.

Details

Nathan Witkin, a researcher at NYU Stern, strongly criticized METR(Model Evaluation and Threat Research)'s 'Long Tasks' benchmark graph as unreliable via 'Transformer' on Substack.

The main points raised are as follows:

  • Data Generalization Error: Data collected from a small number of research colleagues was generalized to represent all of humanity.
  • Baseless Human Baseline: Some human baseline data is based on guesses rather than actual measurements.
  • Flawed Measurement Method: Paying human benchmarkers hourly wages incentivized them to intentionally extend their task time.
  • Sample Bias: The benchmarker sample consists of acquaintances and colleagues of METR staff, lacking representativeness.

Witkin argued that these errors combine to completely undermine the reliability of the results, emphasizing that caution should be exercised when drawing conclusions based on this research.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.