NVIDIA Nemotron 3 Nano Omni: Long-Context Multimodal Intelligence for Documents, Audio, and Video
Key point
NVIDIA has unveiled Nemotron 3 Nano Omni, a multimodal model that handles documents, audio, and video.
Details
NVIDIA Nemotron 3 Nano Omni is an open multimodal model that processes text, images, video, and audio together.
- In document understanding, it recorded OCRBenchV2-En 65.8, MMLongBench-Doc 57.5, and CharXiv reasoning 63.6, targeting long documents, tables, formulas, and multi-page references.
- On the video and audio side, it presented Video-MME 72.2, WorldSense 55.4, DailyOmni 74.1, VoiceBench 89.4, and HF Open ASR 5.95.
- It also achieved OSWorld 47.4 on GUI tasks, and stated that it improved system efficiency by 7.4x and 9.2x respectively over existing open omni models in document and video use cases.
The architecture combines a Nemotron 3 hybrid Mamba-Transformer MoE backbone with a C-RADIOv4-H vision encoder and a Parakeet-TDT-0.6B-v2 audio encoder. For images, it applies dynamic resolution processing with up to 13,312 visual patches, and for video, it reduces tokens using Conv3D tubelet embedding and EVS.
Audio uses 16kHz input, with training inputs supporting up to 1,200 seconds (20 minutes) and the LLM context supporting over 5 hours. SFT was conducted on H100, and the later-stage RL was performed on B200/H100 clusters using NeMo-RL and NeMo Gym. Checkpoints have been released on HuggingFace in BF16, FP8, and NVFP4.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.