AI Briefing
KO

Scenema Audio Model Released

·2026.05.14 06:29

Key point

Scenema Audio has released the weights and inference code for a voice generation model that separates emotion and speaker identity.

Details

Scenema Audio is a zero-shot voice cloning voice generation model that handles emotional expression and speaker identity separately.

  • An emotion prompt specifies the tone of speech, and a reference voice can be added if needed to provide speaker identity.
  • The reference voice serves as the "who" and the prompt serves as the "how," and the model can reportedly generate emotional states for that speaker that it was not trained on.
  • The company has released the model weights and inference code, and has been using it as part of its video production platform.

They emphasize that it sounds more natural and less robotic than autoregressive TTS.

However, due to the nature of diffusion models, some seeds may produce repetition or meaningless utterances, so the workflow assumes selecting and refining candidates after generation.

They also suggested a use case of generating audio first and then fitting video to it via an A2V pipeline (LTX 2.3, Wan 2.6, Seedance 2.0, etc.).

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.