V3 Audio Tags
Key point
ElevenLabs unveiled the Eleven v3 model, which allows fine-grained control over a voice's emotion and delivery through audio tags.
Details
ElevenLabs unveiled a new model, Eleven v3 (alpha version), introducing a feature that controls voice using audio tags instead of text input. Audio Tags is a technology that adjusts non-verbal elements such as emotion, tone, pauses, and speed by entering commands inside square brackets [].
Users can place tags anywhere within the script to direct the voice performance in real time. The main feature categories are as follows.
- Emotions: Set the emotional tone of the voice through tags like
[sad],[angry],[happily] - Delivery direction: Adjust volume and energy through tags like
[whispers],[shouts],[accent] - Human reactions: Achieve natural voice through tags like
[laughs],[clears throat],[sighs]
The new architecture of the v3 model understands text context more deeply, naturally handling emotional cues, tonal shifts, and speaker transitions. This allows interruptions or mood changes in multi-speaker dialogue to be implemented with minimal prompting.
Currently, Professional Voice Clones (PVC) are not fully optimized for v3, so during the research preview stage, it is recommended to use Instant Voice Clone (IVC) or designed voices. Eleven v3 is available through the ElevenLabs UI and the Public API (alpha).
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.