PickNemotron-3-Diarization: Real-time speaker diarization for up to 8 speakers with 80ms latency
nvidia/Nemotron-3-Diarization
About the project
Processes speaker diarization tasks in real time to identify who spoke when. It distinguishes up to 8 speakers and supports both streaming and offline inference, making it suitable for meeting minutes and live broadcast analysis.

Input buffer latency can be reduced to 80ms, enabling operation in environments requiring immediate responses. By applying the Arrival-Order Speaker Cache technique, it retains speaker information from previous chunks, ensuring stable speaker identification even in long audio.
Lightweight execution based on C++ is possible via the NeMo-Speech.cpp runtime, and it can be integrated with ASR pipelines to attach word-level speaker tags. Settings can be flexibly adjusted from offline mode with a 30.4-second buffer to ultra-low latency mode with a 0.32-second buffer.
nvidia/Nemotron-3-Diarization
The original page has no description.
voice-activity-detection
This introduction was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report errors, attribution issues, or removal requests via Contact.