AI Briefing
KO

audio.cpp 0.5 Expands Voice AI

·2026.08.01 09:43

Key point

audio.cpp 0.5 adds expressive TTS, multilingual voice transfer, and AMD GPU support.

Details

audio.cpp 0.5 has added a wide range of new voice AI models and platform features.

  • DramaBox: A prompt-oriented voice acting model based on the LTX-2.3 audio architecture, controlling emotion, tone, laughter, sighs, pauses, transitions, and speaker behavior.
  • Confucius4-TTS: Provides multilingual voice transfer functionality that synthesizes speech in other supported languages based on a reference voice.
  • RVC voice conversion
  • BS-RoFormer vocal separation
  • Added GLM-TTS, Kroko ASR, Parakeet-TDT, Inflect Micro v2, Fun-ASR-Nano

Fun-ASR-Nano was developed by the official FunASR team, and audio.cpp has also been registered on the official FunASR distribution platform.

On the infrastructure side, HIP/ROCm support for AMD GPUs has been introduced at an early stage, and Metal performance on Apple Silicon has also been improved. Real-time PCM input and cleaned-up streaming caption deltas have been added to the server/streaming path, making real-world service integration easier.

Some model integrations and non-CUDA backends still have room for optimization, and the project has presented performance improvements, backend testing, documentation, Web UI, and production deployment feedback as tasks for community contribution.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.