AI Briefing
KOSign in

ElevenLabs Details Audio Tags for Eleven v4 with 50+ Examples

·2026.09.30 21:00

Key point

Audio Tags allow users to direct emotional performance in Eleven v4, which ranks first in the Artificial Analysis Speech Arena with 1319 Elo.

Details

ElevenLabs has released a comprehensive guide to Audio Tags for Eleven v4, its top-ranked Text-to-Speech model. These tags are natural-language cues placed in square brackets, such as [whispers] or [excited], that instruct the model on how to deliver specific lines. The feature allows for granular control over emotional output, pacing, and sound effects, moving beyond simple text-to-speech into directed performance.

Core Functionality and Ranking

Eleven v4 currently holds the top position in the Artificial Analysis Speech Arena with an Elo score of 1319, surpassing competitors like Sonic 3.6. Audio Tags enable users to leverage this high-quality baseline by adding specific performance directions. The tags function by carrying forward until a new tag is introduced, allowing for consistent emotional tone across a sentence or paragraph. Users can combine tags with commas, such as [whispering, fearful], to layer nuances.

Tag Categories and Examples

The guide categorizes tags into several distinct types, each with specific examples:

  • Emotion: Covers a wide range including high energy ([excited], [playful]), heated ([mad], [aggressive]), tense ([anxious], [scared]), low energy ([tired], [bored]), and calm ([peaceful], [thoughtful]).
  • Delivery and Volume: Controls intensity independent of emotion, with tags like [shouts], [softly], and [low, threatening].
  • Pacing: Adjusts speed for suspense or comedy, using tags such as [slowly], [rushed], and [pause].
  • Human-like Reactions: Adds non-verbal sounds like [laughs], [sighs], [gasps], and [crying] to enhance realism.
  • Accent and Character: Allows shifting personas within a single voice, supporting tags like [British accent], [French accent], and [pirate voice].
  • Sound Effects: Introduces non-speech events directly into the generation, such as [thunder rumbling], [footsteps], and [door creaking].

Customization and Best Practices

Users are not limited to a fixed list; they can write custom tags in natural language to describe complex situations or character traits, such as [out of breath after running up the stairs] or [a tired detective who has heard it all before]. The model supports up to 10,000 characters per generation in the Text to Speech app, providing ample space for layered directions.

To achieve optimal results, ElevenLabs recommends listening to the emotion wheel before writing, swapping tags rather than rewriting lines if delivery misses, and using one tag per clause to avoid blurring the performance. Audio Tags are compatible with Eleven v4, Eleven v4 Turbo, and Eleven v3 via the ElevenAPI.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.