Meta Releases Real-Time Multilingual ASR Model 'Muse Voice Transcribe'… Tops Benchmarks with 3.1% WER
Key point
Meta released the real-time streaming ASR model 'Muse Voice Transcribe', achieving a top benchmark ranking with a 3.1% WER.
Details
Meta Superintelligence Labs released the real-time audio recognition model Muse Voice Transcribe. This model supports real-time streaming ASR, speaker diarization, and endpointing, and can accurately separate more than 20 speakers. It ranked first in the streaming speech-to-text and public diarization categories on the Artificial Analysis benchmark.
Performance and Benchmarks
The final transcription Word Error Rate (WER) was 3.1%, lower than competing systems (3.4–4.0%). The average Diarization Error Rate (DER) was 17.5%, superior to existing systems (21.1–28.6%). Notably, it successfully tracked speakers in real time even in complex environments with 8 people speaking simultaneously.
Core Technology: Adaptive Delay
To resolve the trade-off between accuracy and latency, adaptive delay technology was applied. It dynamically adjusts delay time based on word difficulty, optimizing WER rewards and latency rewards through reinforcement learning (RL). This achieved an error rate of approximately 3.0% at the 0.16-second mark, forming a Pareto front.
Multilingual and Code-Switching Support
It natively processes more than 25 languages and perfectly recognizes natural code-switching within a single sentence. Through context biasing, it accurately identifies specific keywords or user contexts such as Meta and Menlo Park. Audio is used only for generating transcriptions and is not stored, ensuring privacy protection.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.