audio.cpp 0.6 Adds New Models
Key point
audio.cpp version 0.6 has been released, supporting five new model families including MiniMax-H3 and a WebUI.
Details
audio.cpp 0.6 has been released, adding five new model families: dots.tts, NeuTTS-2e, MuScriptor, MiniMax-H3, and SenseVoice-Small. This brings support to a total of 49 model families and over 70 model variants.
Key updates include:
- Native WebUI introduction: A web interface for user convenience has been added.
- MiniMax-H3 implementation: Supports versatile applications such as text-to-speech (TTS), voice cloning, and music generation, while strengthening core components for building DiT (Diffusion Transformer) models.
- MiniMax-Music3 (Preview): A music generation model is included as a preview version.
- Experimental video frame generation: Leveraging the characteristic of MiniMax-H3's DiT structure to jointly generate audio and video latents, this feature allows for the experimental generation of RGB frame data.
It has currently been tested in CUDA, Vulkan, and HIP environments. Model VRAM usage and inference speed (RTF) may vary depending on the length of the input audio.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.