Hugging Face Launches Open TTS Leaderboard for Scalable Multilingual and Voice Cloning Evaluation
Key point
The new leaderboard uses objective metrics like WER and RTFx to evaluate open-source TTS models, reducing assessment time from weeks to hours.
Details
Hugging Face has launched the Open TTS Leaderboard to address the fragmentation in evaluating the rapidly growing number of open-source text-to-speech (TTS) models. With over 8K TTS models on the Hub as of September 30, 2026, existing arena-style leaderboards struggle to scale, often underrepresenting open-weights models due to hosting complexities and voter inconsistency.
Objective Metrics and Speed
Unlike human-preference arenas, this leaderboard relies on objective metrics to enable rapid evaluation, cutting the time required to assess a model from a couple of weeks to a couple of hours. The key metrics include:
- Intelligibility: Word and Character Error Rates (WER/CER) measured using Qwen3 ASR.
- Speed: Inverse real-time factor (RTFx) for batched offline inference and time-to-first-audio (TTFA) for streaming latency on H200 GPUs and CPUs.
- Speaker Similarity: Cosine similarity between WavLM speaker embeddings of generated audio and reference clips.
Multilingual and Voice Cloning Support
The leaderboard emphasizes multilingual performance and voice cloning, areas where English-only metrics often fail. Models are ranked on datasets like Seed TTS Eval and CV3 Eval. Top performers in English WER include hexgrad/Kokoro-82M, Supertone/supertonic-3, and fishaudio/s2-pro. For multilingual tasks, k2-fsa/OmniVoice, fishaudio/s2-pro, and FunAudioLLM/Fun-CosyVoice3-0.5B-2512 show strong results. Voice cloning capabilities allow for the comparison of models like bosonai/higgs-tts-3-4b and openbmb/VoxCPM2, which improve when reference audio is provided.
Community and Streaming
The platform includes a "Listen" tab for users to compare generated outputs and vote, aiming to integrate community feedback into future rankings. A dedicated Streaming tab ranks models by TTFA, highlighting kyutai/pocket-tts as a top performer for low-latency streaming on both GPU and CPU. Hugging Face plans to open-source the evaluation scripts to encourage community contributions via GitHub.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.