AI Briefing
KO

NVIDIA unveils Nemotron 3 Nano Omni, integrating vision, audio, and language, boosting AI agent efficiency up to 9x

·2026.04.29 01:00

Key point

NVIDIA has unveiled Nemotron 3 Nano Omni, which integrates vision, audio, and language.

Details

NVIDIA has unveiled Nemotron 3 Nano Omni. It is an open multimodal model that combines vision, audio, and text into one, reducing the latency and context loss that occur when switching between multiple models, thereby improving the response speed and accuracy of AI agents. NVIDIA stated that this model has led 6 leaderboards in complex document intelligence and video/audio understanding.

The core is a 30B-A3B hybrid Mixture-of-Experts (MoE) architecture. It integrates vision and audio encoders to replace separate perception models, and NVIDIA explained that it delivers up to 9x higher throughput than other open omni models that maintain the same level of interaction. It is designed for workflows such as customer support, financial document processing, computer use agents, and audio-video reasoning.

The scope of deployment and customization is also broad. NVIDIA provides open weights, datasets, and training techniques together, and domain-specific optimization and evaluation can be performed with NVIDIA NeMo. The model is available on Hugging Face, OpenRouter, and NIM microservice on build.nvidia.com, as well as through NVIDIA Cloud Partners, inference platforms, and cloud services, and can be deployed with the same flow from DGX Spark and DGX Station to data centers and the cloud.

  • H Company used this model to build a computer use agent that quickly interprets 1920×1080 resolution screen recordings.
  • Aible, ASI, Eka Care, Foxconn, Palantir, Pyler have already adopted it, while Dell Technologies, DocuSign, Infosys, K-Dense, Lila, Oracle, Zefr are evaluating it.
  • As needed, it can be combined with Nemotron 3 Super and Ultra, or other proprietary models, to build sub-agent workflows.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.