MindTopo Reveals the Topological Spatial Reasoning Capabilities of VLMs
Key point
MindTopo measured the gap between VLMs' static topological recognition and interaction planning capabilities.
Details
MindTopo is a benchmark evaluating the 3D topological spatial understanding of multimodal large language models. It focuses on structural relationships such as connectivity, enclosure, order, separation, and knots—relationships that persist even when objects are bent or stretched—rather than precise distances or angles.
Based on topological ability classifications from cognitive science, the benchmark covers the following five domains:
- Continuity: Determining whether paths or objects are unbroken
- Separation: Distinguishing whether adjacent elements form a single structure
- Order: Tracking the arrangement of elements during paths or transformations
- Enclosure: Judging whether boundaries create an inside and an outside
- Knots: Discerning whether ropes are actually tied or connected
Each domain is evaluated through reasoning tasks, which involve judging relationships from static scenes, and planning tasks, which involve creating, maintaining, or removing relationships in simulation environments. For example, models must rotate pipes or rearrange blocks, trap moving agents, or untie ropes without passing them through other lines.
A key feature is the use of a controlled simulator to generate scenes, allowing for precise ground truth and difficulty adjustment. This enables the distinction between cases where visually complex scenes are not recognized and cases where structural relationships are not maintained during changes, even after the scene is understood.
Comparisons across various commercial and open-weight models showed that reasoning performance on static images was consistently higher than interaction planning. Both tasks fell short of human levels, with performance degradation particularly pronounced when relationships needed to be preserved across multiple actions.
Errors primarily occurred during the planning stage. While models sometimes exhibited perceptual errors such as missing walls, openings, or intersections, they also proposed actions that lost the overall structure or violated physical constraints by selecting locally plausible actions even after correctly grasping the scene. This demonstrates that stable decision-making in robotics and interactive environments requires the ability to maintain topological relationships over time.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.