audio.cpp updates: Higgs Audio TTS VRAM reduced by 48%, HTDemucs and PocketTTS speedups, and WebUI history feature
Key point
Higgs Audio TTS now runs in under 6GB VRAM, while HTDemucs sees a 2.21x speedup on CUDA.
Details
Recent updates to audio.cpp introduce significant performance improvements and new features for local audio model inference. The most notable change is a 48% reduction in peak VRAM usage for Higgs Audio TTS, allowing it to run with approximately 6 GB of memory. This optimization was contributed by GitHub user mirek190.
Performance Improvements
The update includes substantial speedups and memory optimizations across several models, maintaining parity and correctness:
- HTDemucs: Achieved a 2.21x speedup on CUDA and 1.95x on Vulkan.
- PocketTTS: Delivered a 2.23x speedup on CPU with a 9% RAM reduction.
- ACE-Step family: Reduced VRAM by 6–7% with speedups of 1.06–1.08x on CUDA and 1.16–1.20x on Vulkan.
- MOSS-TTS v1.5 cloning: Reduced VRAM by 21% with a 1.05x CUDA speedup.
- Echo-TTS (Memory Saver): Reduced VRAM by 20%.
- Qwen3-TTS: Reduced VRAM by 16–20%.
- IndexTTS2 / 2.5: Reduced VRAM by 12%.
New Features and Support
The WebUI now includes an experimental generation history feature, enabling users to revisit previous outputs and restore their settings. The library currently supports over 110 audio model families and 190+ variants, with ongoing improvements to memory efficiency and inference speed across CUDA, Vulkan, Metal, AMD/HIP, and CPU backends. The project is also actively seeking contributors for WebUI frontend development and UI/UX design.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.