AI Briefing
KO

Microsoft Releases Ultra-Fast Speech Transcription Model MAI-Transcribe-2

·2026.09.08 18:30

Key point

Microsoft has released MAI-Transcribe-2, which supports 60 languages and transcribes one hour of audio in just 10 seconds.

1 / 9

Details

Microsoft AI (MAI) has released MAI-Transcribe-2, a speech recognition model that supports 60 languages and processes one hour of audio in 10 seconds. The weights are not released, and the model is available exclusively via a hosted API.

Performance and Benchmarks

Compared to the previous version (MAI-Transcribe-1.5), language support has increased by 17 languages, and Speaker Diarization and word-level timestamp support have been added. It ranked 2nd on the Artificial Analysis leaderboard with a WER (Word Error Rate) of 2.04%, and inference latency was reduced from 20 seconds to 10 seconds. In particular, it shows a processing speed 11.41x faster than GPT-Transcribe and 4.56x faster than Gemini 3.5 Transcribe in terms of speed factor.

Pricing and Deployment

A promotional price of $0.10 per hour of audio is currently applied (regular price undisclosed), and it is deployed as an Azure Speech public preview. However, since no SLA applies, use in production environments is not recommended. The WER based on Korean FLEURS is 3.4%, showing performance similar to Gemini 3.1 Pro (3.3%).

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.