AI Briefing
KO

Gepard 1.0, a Streaming TTS Model for Real-Time Conversation, Released as Open Source

·2026.07.08 01:59

Key point

Gepard 1.0, a 0.6B-parameter streaming TTS model optimized for real-time conversation, has been released as open source.

Details

Gepard 1.0, designed for real-time conversational AI, has been released as open source under the Apache 2.0 license. The model adopts a streaming-first approach that generates audio frame-by-frame as text comes in, without waiting for a sentence to be completed.

Key Technical Specifications and Performance:

  • Model Architecture: A roughly 555M-parameter model based on a Qwen3.5 0.8B backbone (14 layers) and Nemo NanoCodec (FSQ, 22.05kHz).
  • Inference Speed: Achieved a 20x real-time factor (RTF) and ~50ms time-to-first-audio (TTFA) on an RTX 5090.
  • Scalability: Can handle up to 256 parallel sequences on a single RTX Pro 6000 Blackwell (96GB VRAM).
  • Features: Supports zero-shot voice cloning using just a few seconds of reference audio, and supports English, Spanish, Portuguese, and Dutch.

Performance Evaluation: On the Seed-TTS-eval benchmark, the model achieved best-in-class perceptual quality with a NISQA-MOS of 4.25, outperforming existing models such as VoxCPM2 and Fish-S2, and delivered the cleanest audio quality in terms of noise and discontinuity. However, due to streaming optimization, there are some trade-offs in speaker similarity (SIM 0.585) and word error rate (WER 0.036).

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.