Gemini 3.1 Flash TTS - A Next-Generation AI Voice Model That Controls Speech Style Through Natural Language
Key point
Google's new TTS model precisely controls voice style using natural language audio tags.
Details
Google has released Gemini 3.1 Flash TTS in preview. It's a text-to-speech model with significantly improved naturalness, expressiveness, and controllability compared to before, enabling developers, businesses, and general users to build AI voice applications.
It is currently available through the following channels.
- Gemini API and Google AI Studio: for developers
- Vertex AI: for enterprises
- Google Vids: for Workspace users
The core change is audio tags. By embedding natural language instructions directly within text, users can finely adjust voice style, speed, and delivery. In Google AI Studio, this enables more sophisticated voice direction through scene setting, per-character Audio Profiles, Director's Notes, and inline tags.
In terms of performance, it achieved an Artificial Analysis TTS leaderboard Elo of 1,211, and was evaluated for its combination of high-quality voice and low cost. It also supports over 70 languages and has built-in native multi-speaker conversation functionality.
On the security and trust side, SynthID watermarking is applied to all generated audio, making it possible to detect AI-generated voice. This supports content authenticity and helps counter misinformation.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.