AI Briefing
KO

GPT-5.5 Scores Only 10.6% on Vision Benchmark

·2026.07.24 04:20

Key point

On the new ActiveVision benchmark, GPT-5.5 scored 10.6% and Claude Fable 5 scored 3.5%, while human participants achieved 96.1%.

Details

An arXiv paper unveils a new benchmark called ActiveVision, consisting of 17 tasks (3 categories), designed to require models to perform iterative visual perception rather than provide a single static description.

Key results:

  • GPT-5.5 (highest reasoning-effort tier): 10.6%, scoring 0 on 11 of the 17 tasks
  • Claude Fable 5 (ranked #1 on most reasoning/coding leaderboards): 3.5%
  • Average of 3 human participants: 96.1%

More notable than the simple benchmark failure is the structural nature of the failure: models cannot compensate for this limitation even by writing their own code. This suggests a fundamental limitation in frontier vision models' ability to perform iterative, active visual processing.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.