AI Briefing
KO
Pick

Qwen-Robot Suite: A Foundation Model Family for Physical World Intelligence

·2026.06.16 11:00

Key point

The Qwen team has unveiled three robot foundation models that connect vision-language understanding to physical action.

Details

The Qwen family has strong cognitive capabilities for the physical world, but a gap still remains between vision-language understanding and actual physical control. To address this, Qwen-Robot Suite bridges this gap through three core foundation models.

  • Qwen-RobotNav: A gateway for mobility that unifies five navigation task families—instruction following, object navigation, target tracking, and autonomous driving—into a single model.
  • Qwen-RobotManip: A foundation model for interaction that converts heterogeneous robot data into a common canonical space, enabling large-scale cross-embodiment learning.
  • Qwen-RobotWorld: A world model that understands physical dynamics, predicting physical futures across manipulation, driving, and navigation through a natural language action interface.

These models align physical actions across different domains with language, aiming to build an agentic system where general intelligence can translate directly into physical action.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.