Xiaomi-Robotics-1
Key point
Xiaomi has unveiled Xiaomi-Robotics-1, a robot policy model built on 100,000 hours of embodiment-agnostic pretraining.
Details
Xiaomi-Robotics-1 is a policy model that applies the training paradigm of LLMs to robot control. It starts from the problem awareness that the robotics field has not benefited from scaling laws due to data scarcity.
For pretraining, the team leveraged 100,000 hours of UMI (embodiment-independent) data across more than 1,700 scenarios (households, commercial facilities, industrial sites, outdoor environments). They built a pipeline that splits long videos into fixed-length clips and has a VLM automatically language-label the gripper and object state transitions in each clip.
In post-training, they performed fine-tuning along two axes—embodiment alignment and instruction alignment—using a cross-embodiment dataset that includes more than 7,200 hours of real robot data. This covers real-life tasks such as tidying a sofa, organizing a shoe cabinet, and storing kitchenware.
The key finding is that the scaling effect translates into actual robot performance. As pretraining data and model size increase, validation action error decreases, and after post-training, real-robot success rate also continues to rise with no sign of saturation.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.