AI Briefing
KO

Voxtral transcribes at the speed of speech

·2026.02.04 09:00

Key point

Mistral has released Voxtral Transcribe 2, which supports 13 languages, along with a real-time model.

1 / 2

Details

Voxtral Transcribe 2 is two speech-to-text models aimed at high-precision speaker diarization, real-time transcription, and ultra-low latency. The lineup splits into Voxtral Mini Transcribe V2 for batch transcription and Voxtral Realtime for live apps.

Voxtral Realtime uses a streaming architecture that transcribes audio the instant it comes in, rather than chunking and processing an offline model. Latency can be set as low as sub-200ms, and at 2.4-second latency it shows quality suitable for captions, while even at 480ms it claims to keep word error rate loss to around 1-2%.

Language support covers 13 languages: English, Chinese, Hindi, Spanish, Arabic, French, Portuguese, Russian, German, Japanese, Korean, Italian, and Dutch. Designed at 4B parameters, it runs efficiently even on edge devices, and the released weights are distributed under the Apache 2.0 license.

Voxtral Mini Transcribe V2 is a model focused on batch transcription, offering the following features.

  • Speaker diarization: provides speaker labels along with start/end times for who spoke when
  • Context biasing: specify up to 100 words or phrases so names, technical terms, and domain-specific expressions are recognized correctly
  • Word-level timestamps: precise timestamps for each word
  • Noise robustness: strong performance even in environments like factories, call centers, and field recordings
  • Longer audio support: processes up to 3 hours of audio at once

Performance and pricing were also emphasized. Mini Transcribe V2 achieves about 4% word error rate on FLEURS at $0.003/min, which is described as more accurate than GPT-4o mini Transcribe, Gemini 2.5 Flash, Assembly Universal, and Deepgram Nova, and about 3x faster than ElevenLabs Scribe v2 while costing one-fifth as much.

Mistral also unveiled a newly launched audio playground. In Mistral Studio, you can upload audio and immediately test diarization, timestamp granularity, and context bias terms, supporting up to 10 files, 1GB per file, in formats .mp3, .wav, .m4a, .flac, and .ogg.

The use cases are clear. From meeting notes, voice agents, and contact center automation to live captions and regulatory compliance and audit trails, the direction is to make voice workflows that need low latency and speaker diarization cheaper and faster.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.