AI Briefing
KO

Fast Text-to-Audio Generation via Adversarial Post-Training

·2025.05.13 16:24

Key point

ARC post-training generates 12 seconds of 44.1kHz stereo audio in 75ms on an H100.

Details

Text-to-audio systems have improved in performance, but their slow inference speed has become a major problem for creative use due to high latency. To address this, Adversarial Relativistic-Contrastive (ARC) post-training is proposed, and it is presented as the first non-distillation-based adversarial acceleration method applied to diffusion/flow models.

ARC combines two elements:

  • Extending the recent relativistic adversarial formulation to diffusion/flow post-training
  • Newly introducing a contrastive discriminator objective to further improve prompt adherence

Adding several optimizations on top of Stable Audio Open, the resulting model generates about 12 seconds of 44.1kHz stereo audio in about 75ms on an H100. It also achieves around 7 seconds on mobile edge devices, and the authors claim it is the fastest text-to-audio model as of the time of writing.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.