Zyphra Releases Open-Source Real-Time TTS 'ZONOS2'
Key point
Zyphra has launched ZONOS2, an open-source MoE-based TTS model capable of high-fidelity voice cloning and real-time inference.
Details
Zyphra has unveiled ZONOS2, its next-generation real-time text-to-speech (TTS) model. This model is an open-source model provided under the Apache 2.0 License, designed to solve the trade-off between quality and speed.
Key Technical Features:
- Sparse MoE (Mixture of Experts) Architecture: While it has a total of 8B parameters, it maximizes efficiency by using only 900M active parameters during inference.
- High-Fidelity Audio Generation: It generates 44.1 kHz studio-quality audio by predicting Descript Audio Codec (DAC) tokens.
- Language Support and Code-Switching: By directly reading raw UTF-8 bytes without a separate phonemizer, it enhances performance for low-resource languages such as Korean, Chinese, and Japanese, and supports mid-sentence code-switching.
- Zero-shot Cloning: It offers powerful voice cloning performance that captures a speaker's unique characteristics without additional fine-tuning.
ZONOS2 was trained on a massive audio dataset of over 6 million hours, and it demonstrated high expressiveness by recording superior Prosody scores compared to existing major TTS models (Qwen 3 TTS, Cartesia Sonic 3.5, etc.).
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.