Google Releases Gemini 3.8 Flash TTS: Voice Customization via Natural Language Prompts
Key point
It achieved first place overall in the Hume AI benchmark and offers a library of over 2,000 production-ready voices.
Details
Google has released Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, shifting voice generation technology from static presets to a dynamic creative studio. Gemini 3.8 Flash TTS is specialized for deep creative direction and character design, allowing customization of roles, accents, and vocal characteristics in over 100 languages and dialects using only natural language prompts. In contrast, Gemini 3.8 Flash-Lite TTS is optimized for dubbing, audio content creation, and voice agents, targeting bulk processing and cost-effective scaling.
Both models are available in Google AI Studio, Gemini API, Gemini Enterprise, Gemini Notebook, and Google Vids. The voice library provides over 2,000 production-ready voices, including regional variants such as Mexican Spanish, Quebec French, and Scottish English. The Voice Replication feature generates consistent voice profiles from 30-second audio samples, protected by verification of the owner's verbal consent, SynthID watermarking, and C2PA credentials. However, voice replication in AI Studio is restricted in Illinois, Texas, the EEA, the UK, Switzerland, and India.
Directorial features include line-by-line direction via stage directions within scripts, long-form generation that maintains high quality even in hours of continuous audio, and native two-speaker scene staging. Precise control over comedic timing and reaction beats is possible through non-verbal cues like <laughs> and <gasp>, and active listening insertions like |mhm|. The Voice Remixing feature is coming soon.
In benchmark performance, Gemini 3.8 Flash TTS took first place overall (71.4) and first place in accent modeling (60.8) in Hume AI's Voice Design Benchmark. In the Overall Quality Index, Flash TTS ranked first and Flash-Lite TTS ranked second. In Voice Arena's blind human preference evaluation, it ranked among the top competitors in major global languages including Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic, Mexican Spanish, and Hindi. Developers can build high-performance voice generation experiences via the Gemini API on platforms such as Agora, LiveKit, Pipecat, and Vercel, while partners like Figma and HeyGen are integrating the models to accelerate global dubbing and localization.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.