audio.cpp 0.4, Supports 35 TTS Model Families
Key point
audio.cpp 0.4 adds Higgs Audio v3 TTS and Fish Audio S2 Pro among others, now supporting 35 model families with Q8 GGUF improving speed and VRAM.
Details
audio.cpp 0.4 has been released, with high-quality TTS model support and first-class GGUF support as the key changes.
New model support:
- Higgs Audio v3 TTS 4B
- Fish Audio S2 Pro
- Voxtral Realtime ASR
- OuteTTS, VieNeu-TTS-v3 (community models)
The total number of supported model families has grown to 35, and all release models support GGUF.
CUDA Q8 GGUF performance on RTX 5090:
- Higgs Audio TTS: 8.8x–10.1x real-time (after warmup), 8.5x for long-form generation
- Fish Audio S2 Pro: 3.1x–3.4x real-time, 3.3x for long-form
- Voxtral ASR: 15.7x offline, 171ms streaming TTFT
Compared to 16-bit GGUF, Q8 shows up to 1.5x speed improvement and up to 37% reduction in peak VRAM, though results vary by model, so the support matrix remains publicly maintained.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.