AI Briefing
KO

NVIDIA's Real-Time Full-Duplex Voice Model

·2026.08.05 09:00

Key point

NVIDIA released an 11B voice model supporting 450ms response times and tool calling.

Details

NVIDIA NemotronLabs VoiceChat 11B is an end-to-end Full-Duplex voice model that handles speech understanding and generation in a single model via streaming. Unlike the traditional ASR→LLM→TTS architecture, it reduces inter-model connections and API handoffs to target a turn-taking latency of approximately 450ms.

Key features include:

  • Natural turn-taking and real-time speech generation
  • Barge-in, which immediately stops speech when the user interrupts
  • Real-time tool calling that executes while maintaining conversation flow
  • Customized on-hold messages that can be spoken immediately after a tool call

The architecture consists of a Fast Conformer Speech Encoder, a Nemotron Nano v2 9B LLM backbone, and an NVIDIA TTS decoder and codec. The speech signal is converted into audio tokens, after which the LLM predicts text tokens, and the TTS decoder generates speech codes. The tool calling script is predicted on a separate output channel.

This model is introduced as the first publicly available Full-Duplex model to support tool calling, and it ranked 2nd in the open Full-Duplex category on VoiceBench. It is permitted for research purposes only, and usage terms follow the OpenMDW License Agreement 1.1.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.