AI Briefing
KO
Pick

Qwen-RobotManip: Achieving Scaling of Robot Manipulation Foundation Models Through Alignment

·2026.06.16 09:00

Key point

Qwen-RobotManip is a VLA foundation model with outstanding generality, trained on 38,100 hours of large-scale data through data alignment techniques.

Details

Qwen-RobotManip is a general-purpose Vision-Language-Action (VLA) foundation model built on top of Qwen-VL. To address the heterogeneity of existing robot data and the problem of collection costs, it introduces a Three-Dimensional Alignment framework that unifies the Representation, Motion, and Behavior dimensions.

Without using any proprietary data, this model built a pre-training corpus of approximately 38,100 hours using only open-source robot data and human demonstration videos. In particular, through a Human-to-Robot synthesis pipeline that converts 1,933 hours of egocentric human videos into 15 robot embodiments, it generated 24,808 hours of robot demonstration data.

The key achievements are as follows:

  • OOD (Out-of-Distribution) Generalization: It recorded overwhelming performance compared to existing SOTA models on major benchmarks such as LIBERO-Plus (91.4%) and RoboTwin-C2R Hard (69.4%).
  • Real-World Performance: It ranked 1st in the Generalist track of RoboChallenge Table30 v1, demonstrating performance 20% ahead of the 3rd-place model.
  • Cross-embodiment Transfer: It overcomes differences between various robot embodiments and sensors, enabling consistent signal extraction.

As a result, Qwen-RobotManip demonstrates that data diversity and alignment are essential prerequisites for scaling robot foundation models.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.