AI Briefing
KO

PickVibeVoice-ASR-Streaming-7B: Streaming ASR with Real-Time Speaker Diarization Supporting 10 Languages

microsoft/VibeVoice-ASR-Streaming-7B

·2026.09.07 08:29

It transcribes speech in real time, identifying who said what as soon as audio arrives. Unlike traditional ASR models that process sentence by sentence, this streaming approach performs speaker identification without interrupting the flow of conversation.

It supports 10 languages, including Korean, English, Chinese, and Japanese. By registering user-specified proper nouns or technical terms as 'Hotwords', recognition accuracy for those words can be improved, making it advantageous for processing domain-specific data.

Developed by Microsoft Research, this 7B parameter model is released under the MIT License, allowing for commercial use. It is compatible with the Transformers library and is suitable for tasks requiring speaker diarization, such as real-time meeting minutes and customer service analysis.

HuggingFace
HuggingFace model

microsoft/VibeVoice-ASR-Streaming-7B

The original page has no description.

automatic-speech-recognition

This introduction was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report errors, attribution issues, or removal requests via Contact.