AI Briefing
KO
Pick

Qwen-VLA: Going Beyond Understanding the World to Directly Acting in It

·2026.05.29 18:00

Key point

Qwen-VLA is a general-purpose VLA model that goes beyond visual understanding to perform continuous robot actions and trajectory prediction.

Details

Existing multimodal LLMs are skilled at understanding and reasoning about images and video, but they have not reached the stage of Embodied Intelligence, where tasks are directly performed in the physical world. Qwen-VLA is a general-purpose Vision-Language-Action (VLA) model that extends visual perception, language understanding, and spatial reasoning capabilities into continuous Action Generation and trajectory prediction, enabling the model to not just see and think, but to directly act.

Unlike existing approaches specialized for particular tasks, Qwen-VLA provides a unified framework supporting robotic manipulation, Vision-and-Language Navigation (VLN), and a variety of robot platforms. The Qwen multimodal backbone processes visual and language inputs, and the Action Decoder generates continuous actions based on them.

The model's training goes through a sophisticated 4-stage pipeline, from linguistic priors to closed-loop control.

  • Stage I: T2A (Text-to-Action): Learns action structure using only language and embodiment prompts, without images.
  • Stage II: CPT (Continual Pretraining): Jointly trains the VLM and Action Decoder to combine language-action priors with real visual environments.
  • Stage III: SFT (Supervised Fine-Tuning): Fine-tunes using multi-task or real robot data.
  • Stage IV: RL (Reinforcement Learning): Directly optimizes closed-loop task success rate via PPO.

For training data, the model leverages a vast amount of multimodal data including over 10,000 hours of public robot trajectories, more than 8 million synthetic simulation data points, and human egocentric data such as Ego4D, securing generality across diverse environments and tasks.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.