Qwen-RobotWorld: An Infinite World for Embodied Agents
Key point
Qwen-RobotWorld is a next-generation Embodied World Model that uses natural language as an interface to jointly train across diverse robot embodiments and environments.
Details
Existing video generation models possess rich visual priors but lack physical law modeling, while domain-specific models suffer from limited generality. Qwen-RobotWorld bridges this gap by using natural language as a universal action interface.
This model adopts a Dual-Stream Diffusion World Model architecture. An Understanding stream using Qwen2.5-VL as the action encoder and a Generation stream using a video VAE interact through MMDiT (Multimodal Diffusion Transformer), converting complex instructions into precise physical state changes.
Key features are as follows:
- Language-based Unified Action Interface: Standardizes over 20 robot embodiments and 500+ action categories into natural language for unified training.
- Scene2Robot: Provides Human-to-Robot Transfer functionality that retargets human motions to 14 robot embodiments.
- Multi-View Geometric Consistency: Synchronizes 2 to 4 camera streams, including main view and wrist-mounted views, to consistently generate 3D spatial information and object identity.
Through the EWK (Embodied World Knowledge) dataset, the model secured 8.6 million cross-scenario training pairs, enabling physical knowledge transfer across diverse domains such as Manipulation, autonomous driving, and indoor navigation.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.