LingBot-Video Open-Sourced
·2026.07.09 02:58
Key point
LingBot-Video, an action-conditioned world model that predicts robot motion using a Sparse-MoE architecture, has been released.
Details
LingBot-Video is a single-stream Diffusion Transformer model that adopts a DeepSeek-V3-style Sparse MoE architecture. It has 13B total parameters, but uses only 1.4B active parameters during inference.
Key features are as follows:
- RL Post-training: Reinforcement learning was performed using a 6-component reward scheme, including physical-plausibility rewards.
- Action-Video Mode: It predicts robot motion in video form based on the robot's Action and Hand-pose conditions.
- Performance: It achieved the best average performance on RBench, and also demonstrates excellent performance in terms of video generation quality.
The model weights, code, and Diffusers/SGLang stack are currently open-sourced.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.