Tontaube Releases Open TTS Model
Key point
Tontaube has released TontaubeV1, a 2.9B parameter open-source TTS model specialized for long-form generation and low-latency inference.
Details
Tontaube co-founders released TontaubeV1, a 2.9B parameter open-source TTS model focused on long-form generation, expressive speech, and low-latency local inference. The model primarily supports English and German and offers zero-shot voice cloning via reference audio up to 1 minute in length.
Technical Features and Performance
TontaubeV1 uses four separate autoregressive models to generate codec streams from coarse semantic structure to acoustic details, with model sizes progressively decreasing as codebook levels increase. It utilizes character-level text tokenization and shared logical positions between aligned text and audio. A rolling context window enables streaming and virtually unlimited long-form generation, while vLLM-based inference supports batching across requests and acoustic stages.
Measurements on an RTX 5090 achieved approximately 0.08 RTF (Real-Time Factor) for single text inputs and up to 0.02 RTF (approximately 50x real-time speed) with batching. The latency until the first audio encoding is approximately 200ms.
Hardware Requirements and Benchmarks
The current release requires a minimum of 24GB VRAM for low-capacity and balanced profiles, and 32GB VRAM for high-throughput profiles. This is due to the vLLM implementation designed for high concurrency and low-latency serving. Quantized versions for smaller GPUs and compact devices are planned for future release.
In an LLM evaluation based on 400 sentences (audiobook benchmark), TontaubeV1 scored 50.1% against ElevenLabs Flash v2.5 in terms of prosody and was preferred over Fish Audio S2 Pro, Gradium, and Cartesia. Large-scale human listening tests have not yet been conducted but are planned for validation through platforms such as TTS Arena.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.