AI Briefing
KO
Pick

AI Dubbing API

·2026.08.07 05:36

Key point

The AI Dubbing API generates multilingual audio while preserving the original emotion and timing.

Details

AI Dubbing API takes audio or video as input and returns dubbed results in multiple languages, automating speech recognition, translation, speaker voice cloning, synthesis, and sync adjustment. Existing pipelines combining Speech to Text, machine translation, and Text to Speech suffer from issues such as unstable speaker identity, misaligned timing, and monotone emotional expression.

ElevenLabs Dubbing v2 is an audio-to-audio model that directly conditions on the performance and voice of the original recording. By reflecting the original tone, emotion, and delivery style, it generates more natural dubbing than simply reading translated text.

Key features include:

  • Support for 90+ languages
  • Translation and sync adjustment considering original timing
  • Speaker separation and voice preservation
  • Speaker-specific dubbing via voice cloning
  • Creating dubbing projects, adding target languages, and downloading results via REST API
  • Editing source and translated text, and regenerating only changed segments for enterprise customers

Developers can create dubbing projects, add target languages, and download the finished audio or video with just a few API calls.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.