Rhoda AI Demonstrates That Scaling Web Video Pre-training Improves Real-World Robot Task Performance
Key point
Rhoda AI experimentally demonstrated that scaling the size of web video pre-training models and expanding computing resources increases the completion rate of real-world industrial robot tasks.
Details
Rhoda AI used the DVA model to verify whether scaling web-video pre-training leads to improved real-world robot performance. The experimental results showed that increasing the number of parameters and computing resources in the pre-training model consistently improved robot policy performance, a trend that held up to the maximum test size.
Experimental Design and Evaluation Metrics
The experiment followed a method of pre-training on web video followed by post-training on robot data. The quality metric used was DINO FD (the distribution distance of DINOv2 features between held-out web video predicted frames and actual frames), where lower values indicate better prediction performance. Real-world task evaluation included 14-step industrial manipulation tasks such as bearing unpacking and packaging material classification, requiring over 200 hours of evaluation time. The evaluation metric applied was the At-speed completion rate (completion rate based on average cycle time), which is more precise than simple completion rate.
Impact of Model Size and Computing Resources
When scaling the model size from XS to L, the At-speed completion rate increased monotonically from 4% to 85%. Notably, increasing model size contributed more to reducing completion time than to improving accuracy, thereby increasing first-attempt success rates and reducing retries. When fixing the model size at M and increasing pre-training computing resources from 0.08x to 1x, performance improved from 67% to 75%, but performance gains plateaued after 0.37x.
Value of Pre-training in Data-Scarce Scenarios
The effect of pre-training computing resources was most pronounced when post-training data was scarce. When using the full dataset (100%), differences between computing resource levels were negligible, but when reducing the data volume to 25%, a performance gap of 8 percentage points emerged between models with 1x and 0.18x computing resources. This suggests that large-scale web video pre-training provides the greatest advantage when robot demonstration data is limited.
Correlation Between DINO FD and Robot Performance
DINO FD was confirmed as a valid metric for predicting robot task performance, regardless of how model size or computing resources were increased. Models with similar DINO FD scores showed no statistically significant performance differences. However, the relationship between DINO FD and actual performance is merely correlational, with causality unverified, and limitations exist due to the reliance on a single architecture and single task.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.