AI Briefing
KO

Scenema unveils zero-shot voice cloning

·2026.05.14 21:29

Key point

Scenema Audio has released weights for a zero-shot, emotion-decoupled voice generation model.

Details

Scenema Audio has released model weights and inference code.

The key idea is decoupling emotional expression from speaker identity under zero-shot conditions.

  • The reference audio provides who is speaking.
  • The prompt determines how they speak — anger, sadness, excitement, childlike surprise, etc.
  • They explain that even a voice never recorded with that emotion can be generated with the same emotion.

The model is a diffusion model, so it's not as fully stable as traditional TTS.

  • Some seeds produce repetition or nonsensical output.
  • Results vary by seed, making it suited to a generate-then-curate/edit workflow.

As a production workflow, they also proposed audio-first video generation.

  • First generate the voice,
  • then feed it into an A2V pipeline such as LTX 2.3, Wan 2.6, or Seedance 2.0
  • to generate video that syncs with the speech.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.