Qwen3-TTS Update! 49 Voices + 10 Languages + 9 Dialects
Key point
Qwen3-TTS-Flash now supports 49 voices, 10 languages, and 9 dialects.
Details
Qwen3-TTS-Flash has been updated to a flagship text-to-speech model supporting multi-voice, multilingual, and multi-dialect speech synthesis. It aims for natural and expressive speech and is available via the Qwen API.
The biggest change is 49+ high-quality voices. It offers a diverse range of voices spanning gender, age, regional characteristics, and character personality, presenting distinctive role examples such as Momo, Ono Anna, Vivian, Elias, Eldric Sage, and Bunny.
Language and dialect support has also expanded significantly.
- 10 major languages: Chinese, English, German, Italian, Portuguese, Spanish, Japanese, Korean, French, Russian
- 9 dialects: Mandarin, Hokkien, Wu, Cantonese, Sichuanese, Beijing, Nanjing, Tianjin, Shaanxi
- It reported achieving a lower average WER than MiniMax, ElevenLabs, and GPT-4o-Audio-Preview on the MiniMax TTS multilingual test set.
It also explains that compared to the previous version, it adjusts speaking speed and prosody more flexibly to match the text, achieving natural speech that is closer to a human voice. The example samples showcase speaking styles across multiple languages and regional dialects, including Chinese, English, Japanese, and Korean, highlighting emotional expression and intonation differences tailored to real-world user scenarios.
To use it, install the DashScope SDK and call it in the form dashscope.MultiModalConversation.call(...). The model name used is qwen3-tts-flash-2025-11-27, with parameters like voice, language_type, and stream used to specify the voice and language, and the WAV file can be downloaded from the returned audio.url.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.