AI Briefing
KO

Stable Audio 3

·2026.05.21 00:10

Key point

Stable Audio 3 has been released as a model family for variable-length audio generation and editing.

Details

Stable Audio 3 is a fast latent diffusion model series for music and sound generation, presented in three sizes: small / medium / large.

The core features are variable-length generation and audio editing. It's designed to handle everything from sounds a few seconds long to audio spanning several minutes, and it also supports continuation, which extends a short input, and inpainting, which alters only a specific section.

The model runs on a new semantic-acoustic autoencoder. It compresses audio into a smaller latent space to improve diffusion generation efficiency, while reducing quality loss and preserving semantic structure within the latent.

It also applies adversarial post-training to boost both inference speed and quality together. The paper explains that this achieves better fidelity and prompt adherence even with fewer inference steps.

In terms of performance and deployment, the following points are highlighted:

  • Generates music and sound in under 2 seconds on an H200 GPU
  • Generation possible within a few seconds even on a MacBook Pro M4
  • small and medium weights are released, allowing it to run on consumer-grade hardware
  • Training and inference pipelines are also released together

The training data consists of licensed and Creative Commons data.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.