AI Briefing
KO

Prompted TTS

·2026.04.16 02:13

Key point

Google has released Gemini 3.1 Flash TTS, enabling voice style control through prompts.

Details

Google has released Gemini 3.1 Flash TTS. It's available in the standard Gemini API under the model ID gemini-3.1-flash-tts-preview, and the output is provided only as an audio file.

The prompting guide is quite distinctive. Even for synthesizing short sentences, the examples specify lengthy details including character setup, scene description, director's notes, speaking style, pace, and intonation.

Simon Willison tried generating audio with the same example prompt, and noted that changing the setting from being from Brixton to Newcastle or Exeter, Devon didn't significantly change the result. In other words, fine-grained control over regional accents still appears to have limitations.

He also mentioned that he built this experimental UI himself using Gemini 3.1 Pro. The screen shows an API Key input field, a Multi-Speaker (Conversation) mode, speaker name and voice selection, a script input field, and an audio generation button with a download link.

The key point is the emergence of a voice generation model that's closer to prompt-based acting direction rather than simple text-only TTS. However, further verification is needed to see how precisely the actual results reflect these instructions.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.