AI Briefing
KOSign in

Guide to Using Pronunciation Dictionaries and IPA in Eleven v4 TTS

·2026.10.01 21:00

Key point

Eleven v4 supports phoneme rules and inline IPA but does not support SSML phoneme tags.

Details

Eleven v4 offers granular control over text-to-speech pronunciation through pronunciation dictionaries and inline IPA, addressing challenges with brand names, acronyms, and technical terms. While the model handles complex text naturally, users can enforce specific pronunciations for unique vocabulary using alias or phoneme-based rules.

Pronunciation Control Methods

Users can customize speech output using three primary methods in Eleven v4:

  • Inline IPA: Write International Phonetic Alphabet transcriptions directly into scripts between forward slashes (e.g., /ˌkuː.bərˈnɛt.iːz/) for one-off fixes.
  • Phoneme Rules: Define exact phonetic spellings in a pronunciation dictionary for recurring words.
  • Alias Rules: Replace specific text with alternative spellings (e.g., "SaaS" to "sass") before the model speaks.

The model supports up to three pronunciation dictionaries per request via ElevenAPI. Dictionaries can be created in the Text to Speech app, uploaded as .pls files, or managed through the API.

SSML Compatibility and Alternatives

Eleven v4 does not support SSML phoneme tags. Instead, it provides direct replacements for standard SSML functionality:

  • <phoneme> is replaced by inline IPA or dictionary phoneme rules.
  • <sub alias> is replaced by dictionary alias rules.
  • <break> is replaced by Audio Tags like [pause], ellipses, or line breaks.
  • <prosody rate> is replaced by pacing Audio Tags like [slowly].

This approach allows for fluent pronunciation across 90+ languages while maintaining consistent sonic branding.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.