Xiaomi Validates VLA Scaling with 100,000 Hours of Data
Key point
Xiaomi validated VLA scaling using 100,000 hours of manipulation data collected without robots.
Details
The Xiaomi Robotics team released Xiaomi-Robotics-1 (XR-1), a Vision-Language-Action (VLA) model pre-trained on 100,000 hours of real-world manipulation trajectories.
The key innovation is collecting data using a camera-equipped handheld gripper, UMI (Universal Manipulation Interface), instead of robot teleoperation, thereby reducing data bottlenecks associated with the number of robot hardware units.
Language labeling was automated rather than performed manually by humans; trajectories were divided into fixed-length segments, and a VLM was used to describe the state changes in each segment.
The research team conducted pre-training experiments controlling for data volume and model size, then applied the trained checkpoints to real robots to evaluate whether scaling effects translated not only to validation loss but also to real-world robot success rates. The technical report was released on July 16, 2026, and the code and checkpoints were released under the Apache 2.0 license on August 3.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.