AI Briefing
KO

DramaBox, Voice Model Released

·2026.05.14 02:06

Key point

Resemble AI has released DramaBox, a voice synthesis model based on LTX 2.3.

Details

Resemble AI has released DramaBox.

This model is an IC-LoRA tuning layered on top of LTX-2.3 3.3B audio-only, a prompt-based TTS that controls speaker identity, emotion, tone, laughter, sighs, pauses, and transitions using only prompts. Optionally, a 10+ second voice reference can be provided to clone the target timbre, and text embeddings use Gemma 3 12B.

  • A GitHub, Hugging Face model card, and Hugging Face Space were released together.
  • The inference weights are dramabox-dit-v1.safetensors at 6.6GB, and the audio component is 1.9GB.
  • On a warm server, generation time is presented as about 2.5 seconds, with peak VRAM at about 24GB.
  • Outputs are watermarked by default with Resemble Perth's neural network watermark, which can be disabled with --no-watermark.
  • The base model is Lightricks/LTX-2.3, and the license is the LTX-2 Community License.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.