HuggingFace Unveils Compact VLA Model 'SmolVLA'
Key point
SmolVLA, an open-source 450M-scale robotics VLA model that can run on consumer hardware, has been released.
Details
HuggingFace has released SmolVLA, a lightweight Vision-Language-Action (VLA) model designed to lower the barrier to entry for robotics research.
SmolVLA is a model with 450M parameters that runs smoothly on consumer hardware. It was trained using open-source datasets from the Lerobot community, and has demonstrated performance surpassing existing large-scale VLA models and ACT baselines in LIBERO and Meta-World simulations as well as SO100/101 real-world robot environments.
The key technical features are as follows:
- Flow Matching Transformer: Generates precise motions as an Action Expert.
- Inference Efficiency: Applies Visual Token Reduction and Layer Skipping techniques.
- Asynchronous Inference: Improves response speed by 30% and doubles (2x) task throughput.
Along with the model weights, this project encourages the use of affordable open-source hardware, aiming to democratize research into general-purpose robotic agents.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.