ServiceNow releases code-switching speech ASR benchmark
Key point
A new benchmark and dataset have been released to evaluate speech recognition performance on code-switching speech that mixes multiple languages.
Details
More than half of the world's population speaks two or more languages, and code-switching—switching languages mid-sentence—occurs frequently. ServiceNow AI's team built a new benchmark to measure how accurately enterprise voice agents recognize these complex speech patterns.
Key Features and Dataset Composition:
- Language pairs: 4 pairs—Spanish-English, French-English, Canadian French-English, and German-English.
- Scenarios: Reflects real business situations centered on HR (benefits, payroll inquiries) and IT service management (password resets, VPN access, etc.).
- Evaluation metrics: In addition to the simple error rate metric WER (Word Error Rate), the benchmark uses SWER (Semantic WER), which measures meaning preservation, and AER (Answer Error Rate), which indicates downstream task performance.
Benchmark Results: After testing various ASR models, ElevenLabs Scribe V2, Gemini 1.5 Flash, and AssemblyAI Universal 3-Pro recorded top-tier performance across all metrics. The dataset and benchmark tools have been released through AU-Harness.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.