AI Briefing
KO

NVIDIA Unveils Nemotron Nano Omni

·2026.04.29 01:12

Key point

NVIDIA has unveiled the multimodal Nemotron Nano Omni 30B A3B Reasoning.

Details

NVIDIA Nemotron 3 Nano Omni is a multimodal model that processes video, audio, images, and text together.

  • Release date: Released on 2026-04-28 on Build.Nvidia, Hugging Face (BF16/FP8/NVFP4), and NGC.
  • Architecture: A Mamba2-Transformer hybrid MoE structure that combines the Nemotron 3 Nano LLM (30B A3B), the CRADIO v4-H vision encoder, and the Parakeet speech encoder.
  • Specs: About 31B parameters, supporting up to 256k tokens of context, accepting up to 2 minutes of video and up to 1 hour of audio. Language support is English only.
  • Features: Supports text output, JSON format, reasoning output, tool calling, and word-level timestamped transcription.

Key use cases are enterprise workloads such as customer support, media and entertainment analysis, document intelligence, and GUI automation. Commercial use is allowed under the NVIDIA Open Model Agreement.

For the deployment stack, vLLM 0.20.0 along with NeMo, Megatron, NeMo-RL, TensorRT-LLM, llama.cpp, Ollama, and SGLang were presented, along with compatibility guidance for NVIDIA GPUs including A100/H100/H200/B200/L40S/RTX 5090.

Qwen3-VL-30B-A3B-Instruct, Qwen3.5-122B-A10B, Qwen3.5-397B-A17B, Qwen2.5-VL-72B-Instruct, and gpt-oss-120b were used for improvement.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.