From Where Things Are to What They're For: A Spatial-Functional Intelligence Benchmark for Multimodal LLMs
Key point
SFI-Bench is a new video benchmark that evaluates MLLMs' spatial and functional reasoning with over 1,700 questions.
Details
SFI-Bench is a video benchmark built from diverse first-person indoor video scans, comprising over 1,700 questions that jointly test the spatial and functional intelligence of multimodal LLMs. It targets the ability to understand real environments by connecting perception, memory, and reasoning.
There are two evaluation axes.
- Structured Spatial Reasoning: the ability to understand complex indoor layouts and weave together multiple scenes into a coherent spatial representation
- Functional Reasoning: the ability to infer the purpose of objects and their context-specific utility
Questions are composed of tasks such as conditional counting, multi-hop relational reasoning, functional pairing, and knowledge-grounded troubleshooting. In experiments, current MLLMs consistently struggled to combine spatial memory with functional knowledge and external knowledge together, and this benchmark serves as a reference point for moving toward multimodal agents that are more grounded in real environments.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.