AI Briefing
KO

WanSong's Dual-Stem Music Generation

·2026.07.17 19:08

Key point

The Wan Team has released WanSong, a diffusion-based music generation model.

Details

The Wan Team has released a technical report on arXiv for the music generation model WanSong.

Key features:

  • Pure diffusion design: implemented with diffusion alone, without autoregressive methods or multi-stage pipelines
  • Directly outputs 5-minute-long music in a single generation process
  • Dual stems: simultaneously generates separated vocal and background music tracks

Unlike existing music generation models that use a combination of a language model backbone and a diffusion decoder, achieving this level of performance with diffusion alone is notable.

Editing utility: With support for fine-tuning and customization, the generated dual stems allow vocals and instruments to be handled independently in music editing tools. This is especially useful for developers building music editing software and plugins. By providing separated stems rather than a single mixed track, it can serve as a foundation for developing higher-level tools.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.