Scale AI Introduces HSS Benchmark to Test Intuitive Visual Reasoning in Multimodal Models
Key point
Human participants achieve 93.1% accuracy on the new HSS benchmark, while the top-performing model, GPT-6-astra, reaches only 53.6%.
Details
Scale AI researchers introduced Humanity’s Sixth Sense (HSS), a new benchmark designed to evaluate intuitive visual reasoning in multimodal large language models (MLLMs). Existing benchmarks focus on either expert-level analysis or low-level perception, leaving the implicit temporal, spatial, social, and abstract reasoning humans perform at a glance largely untested.
The HSS benchmark spans diverse image and video inputs organized under a structured taxonomy. Each item is paired with human-written prompts that probe the implicit structures people infer from a single glance or a few seconds of video, such as determining if a vehicle fits between parked cars or identifying authority figures in a room.
Performance Gap Between Humans and Models
Frontier MLLMs significantly underperform compared to human participants on intuitive visual tasks:
- Human Performance: Participants achieved 93.1% accuracy.
- Model Performance: The strongest model, GPT-6-astra, reached only 53.6% accuracy even at maximum reasoning effort.
Despite excelling in complex tasks requiring advanced perception and knowledge, current models struggle with these intuitive visual reasoning capabilities. The researchers also explored an agentic setup applying dynamic visual manipulation to HSS, which narrowed but did not close the performance gap.
HSS establishes intuitive visual reasoning as a measurable axis, highlighting a capability that scaling on current benchmarks has left behind.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.