AI Briefing
KO

Scenema Audio Now Supports 8GB VRAM

·2026.08.06 03:27

Key point

Scenema Audio has been released as a ComfyUI native node that runs on 8GB VRAM.

Details

Scenema Audio has been released as a native custom node for ComfyUI. By quantizing existing API- and Docker-stack-based models, it can run with a minimum of 8GB VRAM.

It offers expressive TTS that allows specifying tone and emotion via text, and also supports zero-shot voice cloning using reference audio.

  • Inline directives such as [he laughs softly] and [voice cracks] are executed at the corresponding utterance position
  • Provides 12 preset voices with different intonation, age, and emotional expression
  • Generation speed is up to 2x real-time
  • Requires downloading approximately 30GB of model weights on first run
  • Uses the gated Hugging Face model Gemma 3 12B as the text encoder

The node can be installed by searching in ComfyUI Manager or directly from GitHub, and official workflows are automatically added to the Workflows sidebar in ComfyUI. Since it is a diffusion-based model, repetitions or broken audio may occur with certain seeds, so the workflow assumes generating multiple results for selection and editing.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.