AI Briefing
KO

4B VLA Model 'Wall-OSS-0.5' Released

·2026.05.29 01:37

Key point

X Square Robot has released Wall-OSS-0.5, a 4B-scale VLA model that includes open-source training code.

Details

Wall-OSS-0.5 is a 4B-scale VLA (Vision-Language-Action) model that combines a 3B VLM backbone with action experts in a Mixture-of-Transformers structure.

Key performance metrics are as follows:

  • Zero-shot performance: Evaluated on a set of 17 real-robot tasks, achieving a progress score of 80 or higher on 4 tasks, including deformable tasks such as 'Rope Tightening'.
  • Fine-tuning performance: Achieved an average progress score of 60.5 when fine-tuned on a set of 15 tasks, a 17.5pp improvement over pi0.5.
  • Embodied grounding: Improved related performance by 21.8pp while maintaining general VLM capabilities.

Key technical elements include the Gradient Bridge concept, in which the CE (Cross-Entropy) of action tokens provides the dominant gradient to the VLM backbone, and the Vision-Aligned RVQ Tokenizer, which semantically grounds action tokens. It also introduces a decentralized DMuon optimizer with reduced overhead.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.