Wan-Dancer: Generating Dance Videos Over 1 Minute Long
Key point
Wan-Dancer, a model that generates high-quality dance videos over 1 minute long in sync with music rhythm, has been released.
Details
A new hierarchical framework called Wan-Dancer has been announced to solve the temporal drift and motion repetition problems that existing Diffusion Models experienced when generating videos longer than 20 seconds.
This model separates the process into Global keyframe planning and Local temporal refinement stages, maintaining long-term consistency while preserving the context of the entire song.
The key technical features are as follows:
- Time-mapped RoPE embeddings: Achieves precise alignment between music and motion through dynamic frame rate adaptation
- Optical-flow-based loss: Applies an optical flow-based loss function to enhance motion continuity
- Motion-speed control: A speed control feature that maintains high-resolution detail even during fast movements
Experimental results show that it generates stable videos over 1 minute long at 720p/30fps resolution, overcoming the temporal limitations of existing models. The model weights and inference code are currently available via GitHub and HuggingFace.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.