AI Briefing
KOSign in

OpenWAM: An Open Framework for Composable World-Action Models

·2026.10.06 09:00

Key point

OpenWAM is an open framework for composable world-action models that achieves 98.6% success on the LIBERO benchmark using a Mixture-of-Transformers architecture trained on 14.64k hours of robot video data.

1 / 4

Details

Researchers have introduced OpenWAM, an open framework for composable world-action models (WAMs) designed to decouple video backbones, interaction structures, and supervision methods. By providing a shared causal robot-video foundation and configurable video-action interaction programs, OpenWAM addresses the difficulty of comparing design choices in existing WAMs that simultaneously alter multiple architectural components.

Architecture and Pretraining

The framework builds upon the Wan2.2-5B model, performing causal robot-video pretraining on approximately 3.34 million trajectories totaling 14.64k hours of source video. This dataset includes real and synthetic robot manipulation, human-guided manipulation, and human interaction from sources like Open X-Embodiment, AgiBot World Beta, and InternData-A1. Training was conducted on 32 NVIDIA B200 GPUs for 14 days.

Key architectural features include:

  • Mixture-of-Transformers (MoT): Combines a 5B video expert and a 2B action expert, initialized from width-adapted copies of pretrained video layers.
  • Causal Attention: Generates video in chunks, where each chunk references only current/past observations and previous chunks.
  • Four Interaction Programs: Supports Video-then-action (VTA), Action-then-video (ATV), Joint, and Decoupled generation orders using latent flow matching.

Performance and Benchmarks

OpenWAM demonstrates strong performance in both simulation and real-world settings:

  • LIBERO Benchmark: The OpenWAM-VTA variant achieved a mean success rate of 98.6% (Object: 99.4%, Goal: 98.4%, Spatial: 98.6%, Long: 97.8%). This outperforms baselines such as LingBot-VA (98.5%) and Motus (97.7%).
  • Real-World Manipulation: On bimanual tasks using two Franka Research 3 arms, OpenWAM-VTA achieved a mean success rate of 92.1% across toasting bread, solving a Rubik's cube final layer, and sorting cups.
  • Ablation Results: Robot-video pretraining improved LIBERO-Long success by 29.4 points for VTA and 34.4 points for Joint compared to original Wan2.2 initialization. The MoT architecture provided a 5.0-point improvement over shared DiT for VTA.

Counterfactual Data and Dynamics Models

The study introduces LIBERO-Long-CF, a dataset of 32,000 counterfactual segments generated by simulating alternative action sequences, including failures. This data significantly enhances the reusability of independently trained dynamics components:

  • Inverse Dynamics Model (IDM): Using counterfactual data, the local-context IDM achieved 84.0% mean success on held-out LIBERO-90 tasks, compared to 21.5% with demonstration-only data.
  • Forward Dynamics Model (FDM): Counterfactual supervision reduced RGB prediction error by 34.5% and improved the ability to identify correct future outcomes from 16 alternatives from 21.1% to 71.3%.

Unified vs. Specialized Models

A unified OpenWAM checkpoint supporting policy, IDM, and FDM objectives achieved 92.8% success on LIBERO-Long, surpassing the released UVA system (88.0%). However, specialized models still offer higher task success and more accurate counterfactual predictions, suggesting a trade-off between composability and peak performance.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.