AI Briefing
KO

Breeze-TTS-2: First audio in 40ms on H100, outperforming closed-source TTS

BreezeBlue/Breeze-TTS-2

·2026.08.28 02:10

Breeze TTS 2 is an open-weight text-to-speech model designed for real-time interaction. It ranks first among open-weight models on the Artificial Analysis TTS leaderboard, demonstrating performance that surpasses some closed frontier systems. A single model naturally generates speech in English and Chinese.

It supports Voice Design, which creates specific voices using only natural language descriptions without reference audio, and Voice Direction, which directs tone, emotion, and speed while preserving the existing voice. By simply inserting notations such as (laugh) or [笑] in the text, it automatically inserts expressive vocal events like laughter, coughing, and throat clearing.

In an NVIDIA H100 environment, using a warmed-up fast path results in a Time to First Audio (TTFA) of less than 40ms. It achieves a Real-Time Factor (RTF) of 0.32, generating audio at approximately 3.1x speed. During inference, GPU memory usage is approximately 7.7 GiB, with a 12GB GPU being the minimum recommended specification.

The source code is licensed under Apache 2.0, but model weights, derivative models, and self-hosted outputs are limited to research and non-commercial use. Commercial use requires written permission from RESONIA, INC. 24kHz PCM data can be received via a real-time streaming API.

HuggingFace
HuggingFace model

BreezeBlue/Breeze-TTS-2

The original page has no description.

text-to-speech

This introduction was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report errors, attribution issues, or removal requests via Contact.