AI Briefing
KO
Pick

NVIDIA Unveils Nemotron 3.5 ASR Supporting 40 Languages

·2026.06.04 21:59

Key point

NVIDIA has unveiled Nemotron 3.5 ASR, a roughly 600-million-parameter model that supports 40 languages in real time.

Details

NVIDIA has unveiled Nemotron 3.5 ASR, a multilingual speech recognition (ASR) model capable of real-time streaming. This model supports 40 language locales with a single checkpoint, and has built-in punctuation and capitalization capabilities.

Key features are as follows:

  • Efficient Streaming: Using a Cache-Aware FastConformer-RNNT architecture, it processes audio frames only once without redundant computation, dramatically reducing latency.
  • Low Latency: In the Artificial Analysis benchmark, it ranked 2nd in latency among streaming ASR models, demonstrating a balance of high accuracy and low latency.
  • Multilingual Support and Auto-Detection: It supports 40 languages including Korean, English, Japanese, and Chinese, and can be configured to either directly specify the input language or have the model automatically detect it.
  • Open Weights and Fine-Tuning: Provided as open weights via Hugging Face, allowing users to optimize it through fine-tuning for specific languages, domains, or accents.

This model generates text at a level ready for immediate service application without a separate post-processing step, and can be deployed on one's own infrastructure without API dependency.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.