SANA-WM, an Open-Source 2.6B-Parameter World Model for 1-Minute 720p Video
Key point
NVIDIA has unveiled SANA-WM, which generates 1-minute 720p video with 6-DoF camera trajectories.
Details
NVIDIA has unveiled SANA-WM. It is a 2.6B-parameter world model that takes a single image and a 6-DoF camera trajectory as input and generates controllable 720p, 1-minute-long video on a single GPU. The page provides Paper, Code, and Models soon links.
The core architecture is a Hybrid Linear Diffusion Transformer.
- It combines frame-wise Gated DeltaNet with periodic softmax to maintain consistency over long rollouts.
- A coarse global pose branch and a fine pixel-aligned geometry branch improve camera path following.
- The all-softmax variant is described as running into OOM during 60-second generation.
Training used about 213K publicly available videos with meter-scale 6-DoF pose supervision.
- Full training took 15 days on 64 H100s.
- Inference is possible on a single H100.
A 17B long-video refiner is attached for quality improvement.
- It refines the texture, motion, and later-segment quality of the stage-1 output.
- The distilled variant denoises a 60-second 720p clip in 34 seconds using NVFP4 on 1 RTX 5090.
On benchmarks, it reportedly showed higher action-follow accuracy than existing open-source baselines, and 36x higher throughput at comparable visual quality.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.