ElevenLabs' New STT Model, Scribe
Key point
ElevenLabs has unveiled Scribe, an STT model supporting 99 languages with industry-leading accuracy.
Details
ElevenLabs has launched its first Speech to Text (STT) model, Scribe. Scribe supports 99 languages and offers the following features:
- Word-level timestamps and speaker diarization
- Audio event tagging (e.g., laughter, etc.)
- Seamless integration via structured JSON responses
Scribe has demonstrated performance surpassing Gemini 2.0 Flash, Whisper Large V3, and Deepgram Nova-3 on the FLEURS and Common Voice benchmarks. In particular, it recorded the industry's lowest word error rate (WER) in 97 languages, including Italian (98.7%) and English (96.7%).
It has also significantly improved accuracy for underserved languages such as Serbian, Cantonese, and Malayalam—where existing models had error rates exceeding 40%—enhancing the versatility of ASR (Automatic Speech Recognition).
Developers can integrate it immediately via the Speech to Text API, while creators and businesses can use it by directly uploading files through the ElevenLabs dashboard. A low-latency version for real-time applications is also set to launch soon.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.