ASR Models Memorize Benchmark Answers
Key point
A study quantifies the phenomenon of ASR models inflating scores by learning benchmark patterns rather than actual speech.
Details
While public Automatic Speech Recognition (ASR) benchmark scores are reported to have reached human-level performance, this may be because models are optimized for the test itself (benchmaxxing) rather than demonstrating actual task capability. A recent study evaluated 11 open-source ASR models and confirmed a tendency to reproduce benchmark transcripts from the VoxPopuli and LibriSpeech datasets verbatim.
Notably, some models were found to ignore the audio content and output the benchmark's 'correct answer', even when the audio contradicted the benchmark's reference transcript or when relevant words were silenced. This is analyzed as models identifying specific benchmark data through subtle acoustic cues in the speech and generating corresponding outputs.
This phenomenon leads to an overestimation of models' actual speech recognition capabilities, suggesting that benchmark scores alone are insufficient for judging real-world performance. To measure this, the research team proposed three quantification methods, including a reference disagreement test.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.