Fast Text-to-Audio Generation via Adversarial Post-Training
Key point
ARC post-training generates 12 seconds of 44.1kHz stereo audio in 75ms on an H100.
Details
Text-to-audio systems have improved in performance, but their slow inference speed has become a major problem for creative use due to high latency. To address this, Adversarial Relativistic-Contrastive (ARC) post-training is proposed, and it is presented as the first non-distillation-based adversarial acceleration method applied to diffusion/flow models.
ARC combines two elements:
- Extending the recent relativistic adversarial formulation to diffusion/flow post-training
- Newly introducing a contrastive discriminator objective to further improve prompt adherence
Adding several optimizations on top of Stable Audio Open, the resulting model generates about 12 seconds of 44.1kHz stereo audio in about 75ms on an H100. It also achieves around 7 seconds on mobile edge devices, and the authors claim it is the fastest text-to-audio model as of the time of writing.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.