AI Briefing
KO

Nvidia Unveils Nemotron 3 Nano Omni

·2026.04.29 06:53

Key point

Nvidia has unveiled Nemotron 3 Nano Omni, an open-weight multimodal model.

Details

Nvidia has unveiled Nemotron 3 Nano Omni. It's an MoE-based open-weight multimodal model that uses only 3B active out of 30B parameters, combining vision, audio, and language in a single architecture, targeting edge and single-GPU environments.

Nvidia claims 9x throughput, 2.9x faster single-stream inference, and about 9x higher video reasoning system capacity compared to open multimodal models of similar scale. It also stated it leads on 6 benchmarks related to document, video, and audio understanding.

The architecture is a hybrid Mamba-Transformer.

  • 23 Mamba-2 layers
  • 23 MoE layers (128 experts, 6 routed per token + shared expert)
  • 6 grouped-query attention layers
  • Image encoder C-RADIOv4-H, audio encoder Parakeet-TDT-0.6B-v2, video via 3D convolution
  • The base text model was pretrained on 25T tokens, and it supports context up to 256,000 tokens.

Commercial use is available under the Open Model Agreement, and it can be deployed via Hugging Face, NIM microservice, Amazon SageMaker JumpStart, OpenRouter, as well as vLLM, SGLang, Ollama, llama.cpp, TensorRT-LLM. Foxconn, Palantir, Aible, ASI, Eka Care, and H Company have adopted it, while Dell, DocuSign, Infosys, Oracle, and Zefr are evaluating it.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.