Microsoft tests new MAI Realtime voice model
Key point
Microsoft is testing MAI Realtime, a full-duplex real-time voice model, in MAI Playground.
Details
Microsoft is testing MAI Realtime, which appears to be its first in-house real-time voice model, with select partners in MAI Playground. The model uses a full-duplex approach that listens to the user while responding at the same time, competing with OpenAI's GPT Live 1 or Sesame.
It currently offers two voices, Victoria and Grant, and supports more natural responses than Copilot's voice mode. It supports 17 languages including English, German, Spanish, French, Italian, Portuguese, Japanese, Korean, and Chinese, and automatically adapts when the language changes mid-conversation.
The turn-taking method for conversations can be set to one of the following two options.
- Switchboard mode: uses the MAI-Ears endpointer with inline control tokens.
- Deterministic mode: combines silence-based endpointing with the Whisper semantic endpointer.
The actual difference between the two methods is not large, but interruption handling is reportedly natural and response latency is also short. It does not generate singing or non-speech sound effects, focusing more on conversational voice systems than general-purpose audio generators.
The debug panel in MAI Playground lets you check real-time latency figures, the model's processing steps, and its internal thought process. As access expands going forward, a feature for Playground users to share samples is also expected to be provided.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.