ElevenLabs General Availability of Medical-Specific Speech Recognition Model 'Scribe v2 Medical'
Key point
ElevenLabs released Scribe v2 Medical, a speech recognition model specialized for medical audio, improving clinical audio WER by 18%.
Details
ElevenLabs has made its speech recognition model specialized for clinical audio, Scribe v2 Medical, generally available (GA) via ElevenAPI. This model recorded the lowest Word Error Rate (WER) on medical ASR benchmarks, reducing clinical audio WER by 18% compared to the existing Scribe v2.
Benchmark Performance and Accuracy
On the Eka Medical ASR benchmark (3,619 samples), Scribe v2 Medical recorded a WER of 7.0%, an improvement of 1.6pp over the Base model (8.6%). Statistical significance was verified through 10,000 reconstruction tests. On the Omi Health benchmark (1,513 clips), it achieved the lowest overall WER of 5.88% among 30 systems, and dosage recognition accuracy improved from 79.8% to 86.2%.
In standalone evaluation of medical terms, errors decreased by 14% (11.1% → 9.5%). Specifically, WER dropped to 7.5% for term recognition within sentences, but isolated single-word recognition without context remains a difficult challenge with a WER of 14.3%. To address this, a keyterm prompting feature is provided to emphasize drug names and other terms.
Non-Medical Performance Retention and Security
It was verified that medical-specific fine-tuning does not degrade performance in non-medical domains. On daily speech samples based on CommonVoice, it maintained the same 5.3% WER as the Base model. Additionally, for enterprise customers who have signed a BAA and enabled Zero Retention Mode (ZRM) for HIPAA compliance, audio and text data are immediately deleted. Users can specify the model ID scribe_v2_medical to call it via the API.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.