AI Briefing
KO

Qwen3-TTS Evolves: Voice Cloning and Voice Design

·2025.12.23 01:00

Key point

Qwen3-TTS now supports natural language voice design and 3-second voice cloning.

Details

The Qwen3-TTS family has released two new models. Qwen3-TTS-VD-Flash is a voice design model that allows fine-grained control of timbre, prosody, emotion, and persona using only natural language instructions, while Qwen3-TTS-VC-Flash is a voice cloning model that supports 3-second voice cloning.

  • VD-Flash is designed to control not just "what to say" but also "how to say it."
  • On InstructTTS-Eval, it comprehensively outperformed GPT-4o-mini-tts and Mimo-audio-7b-instruct, and in role-play tests it was reported to be superior to Gemini-2.5-pro-preview-tts.
  • VC-Flash supports cloned-voice-based synthesis in 10 major languages, including Chinese, English, German, Italian, Portuguese, Spanish, Japanese, Korean, French, and Russian.
  • On the MiniMax TTS Multilingual Test Set, it is stated to achieve better average WER than MiniMax, ElevenLabs, and GPT-4o-Audio-Preview.

Both models aim for highly expressive, human-like voices. They automatically adjust intonation and rhythm to match the meaning of the input text, and handle complex sentence structures or unstructured text stably.

Additionally, users can save generated voices for repeated calls, making them usable for multi-speaker, multi-turn, long-form conversations. The article also includes example code for calling qwen-voice-design via the Qwen API to create voice profiles.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.