AI Briefing
KO

Qwen3-TTS Family Released as Open Source: Voice Design, Cloning, and Generation

·2026.01.22 01:00

Key point

Qwen3-TTS has been released with a 1.7B and 0.6B lineup, supporting voice design, voice cloning, and streaming generation.

Details

The Qwen3-TTS family has been released as open source. It combines Voice Design, Voice Clone, high-quality voice generation, and natural language-based voice control into a single series, and has been released in two sizes, 1.7B and 0.6B.

The core is a multi-codebook speech encoder based on Qwen3-TTS-Tokenizer-12Hz and a Dual-Track hybrid streaming structure. Through this, the speech signal is efficiently compressed and represented, while the non-DiT lightweight structure enables high-speed, high-fidelity reconstruction, with the first audio packet output right after a single character input and end-to-end latency reduced to as low as 97ms.

The model lineup is divided by function.

  • Qwen3-TTS-12Hz-1.7B-VoiceDesign: Designs a voice timbre from a user description
  • Qwen3-TTS-12Hz-1.7B-CustomVoice: Controls 9 premium timbres via instructions
  • Qwen3-TTS-12Hz-1.7B-Base: A base for fast voice clone and fine-tuning using 3 seconds of speech
  • Qwen3-TTS-12Hz-0.6B-CustomVoice / Base: A lighter, balanced lineup

The supported languages are 10 major languages: Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, and Italian, and various dialects are also included. Timbre, emotion, and prosody can be adjusted via natural language instructions, and tone and rhythm are adaptively changed to reflect the text's meaning.

The evaluation results are also strong. In voice design, it outperformed the closed-source MiniMax-Voice-Design model on InstructTTS-Eval in both instruction-following and expressiveness, and in voice control it recorded an average WER of 2.34% and style control of 75.4%. In 10-minute continuous synthesis, it showed WER 2.36% for Chinese and 2.81% for English, and in voice clone, it recorded an average WER of 1.835% across 10 languages and speaker similarity of 0.789, outperforming MiniMax and ElevenLabs.

The tokenizer also showed strong performance on its own. On the LibriSpeech test-clean benchmark, it recorded PESQ 3.21/3.68, STOI 0.96, UTMOS 4.16, and speaker similarity of 0.95, presenting SOTA-level speech reconstruction quality and speaker information preservation.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.