Wispr Unveils 'Canto', a Real-Time Transcription-Specialized Voice Model Achieving Lower WER Than Competitors via GRPO Training
Key point
Wispr Advanced Interfaces Lab released 'Canto', a voice model that achieves low Word Error Rate (WER) in real-time transcription environments by applying GRPO reinforcement learning.
Details
Wispr Advanced Interfaces Lab has released 'Canto', a new voice model designed for real-time dictation. The model was pre-trained on millions of hours of audio data and then trained using a combination of Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL).
Performance and Evaluation
In benchmarks including real transcription data and challenging environments such as noise, whispering, and distant speech, Canto recorded the lowest Word Error Rate (WER) compared to competing models like Google, OpenAI, AssemblyAI, and Deepgram. In the overall challenging set, the large multimodal model Gemini 3.1 Pro took first place; however, Gemini 3.1 Pro is not suitable for real-time low-latency applications. Among real-time transcription models, Canto achieved the lowest WER. Notably, Canto tied for the lowest error rate in low-volume speech and short transcription sentences.
GRPO-Based Training Strategy
Canto utilized Group Relative Policy Optimization (GRPO) to improve training efficiency. It generates multiple transcription candidates (rollouts) for the same audio input, calculates rewards by comparing them with reference data, and updates the model based on relative advantage. This approach reduces recognition errors in difficult recording conditions and establishes a learning environment for specific failure modes through an RL method that evaluates the quality of completed transcriptions.
Future Plans
Wispr is developing next-generation models with Canto as the first step, leveraging infrastructure to solve open research problems such as context reliability, personalization, and speaker diarization.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.