AI Briefing
KO

audio.cpp Expands as a GGML-Based Audio Generation Model

·2026.07.03 12:12

Key point

audio.cpp, a C++/GGML-based framework, has undergone a major update by adding music and SFX generation as well as source separation features.

Details

audio.cpp, a native C++/GGML framework, has announced a major expansion including music and sound effect (SFX) generation, and Source Separation features. Through this update, the framework has evolved beyond voice (TTS), ASR/VAD, and voice conversion into a unified framework supporting audio generation as a whole.

Key Updated Models:

  • ACE-Step 1.5 (Turbo/Base): Supports high-speed music generation
  • HeartMuLa: Capable of generating audio up to 10 minutes long
  • Stable Audio 3 (Small/Medium): Music and SFX generation
  • Mel-Band RoFormer and HTDemucs: Source separation features

Performance and Features:

  • Based on ACE-Step Turbo, it demonstrated efficiency by recording an execution speed about 1.4x faster than Python (RTF 0.100).
  • It supports a mem_saver mode for server environments and long-running execution, allowing VRAM usage to be lowered after execution.
  • The framework's current completeness level is about 75%, and backend performance optimization is planned going forward.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.