AI Briefing
KO

SANA-WM, a 2.6B open-source world model for 1-minute 720p video

·2026.05.16 21:06

Key point

NVIDIA has released a 2.6B open-source world model for 1-minute 720p video.

Details

SANA-WM is a 2.6B open-source world model released by NVIDIA that generates 60-second, 720p controllable video from just a single image and a camera trajectory, based on a Hybrid Linear Diffusion Transformer architecture.

The core structure consists of four elements.

  • Hybrid Linear Attention: combines frame-wise Gated DeltaNet with periodic softmax to maintain long context.
  • Dual-Branch Camera Control: follows 6-DoF paths with high precision using a coarse global pose branch and a fine pixel-aligned geometric branch.
  • Two-Stage Generation Pipeline: after a stage-1 rollout, a 17B long-video refiner enhances texture, motion, and later-segment consistency.
  • Robust Annotation Pipeline: extracts metric-scale 6-DoF pose from public videos to produce well-aligned action labels.

Training used approximately 213K public video clips, completed in 15 days on 64 H100s. The original model generates a 60-second 720p clip on a single H100, and the distilled variant denoises a clip of the same length in 34 seconds on an RTX 5090 using NVFP4 quantization.

On its own one-minute world-model benchmark, it reportedly shows higher action-following accuracy than existing open-source baselines, while maintaining visual quality on par with LingBot-World and HY-WorldPlay, and achieving 36x higher throughput.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.