Character-Level TTS Model Released
Key point
TontaubeV1, a 2.9B parameter open-weight TTS model based on character-level tokenization and DualCodec, has been released.
Details
TontaubeV1 is an 2.9B parameter open-weight text-to-speech (TTS) model released with optimizations for long-form generation and low-latency local inference. The model primarily supports English and German, offering zero-shot voice cloning via reference audio of up to 1 minute.
Character-Level Tokenization Applied
While most recent LLM-based TTS models use the BPE tokenizer of the backbone model, TontaubeV1 is based on the Qwen3-1.7B checkpoint but adopts a method that forces speech text to be tokenized as individual character sequences. The developers stated that this approach simplifies character-to-sound mapping while maintaining language understanding capabilities, thereby reducing out-of-distribution errors caused by rare token sequences.
Training Data and Architecture
The model is built on the DualCodec multi-codebook discrete audio codec and was trained on approximately 200,000 hours of audio data across 7 languages. Currently, testing and validation have been primarily conducted in English and German environments.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.