AI Briefing
KO

ZDTaichu5.0-9B: Combining Spatial Reasoning and Agent Capabilities in a 9B-Class Multimodal Model

TaichuAI/ZDTaichu5.0-9B

·2026.09.17 11:43

It combines the Qwen3.5-9B language backbone with the C-RADIOv4-H vision encoder to process text, images, and video simultaneously. While maintaining general visual understanding capabilities, it supports spatial recognition, 3D scene interpretation, and multi-image and video analysis.

It demonstrates top-tier spatial reasoning performance among 10B-class general-purpose VLMs on benchmarks such as SparBench and ViewSpatial. With scores of 48 on ERQA and 56 on RoboSpatial, it possesses the scene understanding and planning capabilities required for robotics and embodied AI research.

Beyond simple visual recognition, it enables multi-step tool use and agent task execution. It handles complex vision-language tasks such as document, chart, OCR, and visual math problem solving, accepting various inputs without resolution constraints.

Supporting English and Chinese, it is suitable for researchers seeking to acquire the spatial intelligence necessary for embodied AI or robot control logic development. Its key feature is the addition of specialized spatial and agent capabilities without sacrificing general-purpose versatility.

HuggingFace
HuggingFace model

TaichuAI/ZDTaichu5.0-9B

The original page has no description.

image-text-to-text

This introduction was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report errors, attribution issues, or removal requests via Contact.