AI Briefing
KO

NVIDIA Unveils Nemotron-3 Puzzle, Maximizing Efficiency

·2026.07.07 20:32

Key point

NVIDIA has released Nemotron-3 Puzzle-75B, an MoE-based model that dramatically improves inference efficiency.

Details

NVIDIA has unveiled Nemotron-Labs-3-Puzzle-75B-A9B, a model that maximizes inference efficiency by applying the Iterative Puzzle compression framework. This model was developed based on the existing Nemotron-3-Super-120B-A12B.

Key features are as follows:

  • Hybrid MoE Architecture: It adopts a structure combining Mamba, MoE, and Attention layers.
  • Compression and Efficiency: Total parameters were reduced from 120.7B to 75.3B, and active parameters from 12.8B to 9.3B, boosting inference performance.
  • Performance Improvement: Server throughput on a single 8×B200 node was improved by approximately 2x, and 1M-token concurrent processing capacity on a single H100 was expanded from 1 to 8.
  • Multilingual and Capabilities: It supports English, French, German, Italian, Japanese, Spanish, and Chinese, while maintaining strong performance in reasoning, coding, and agentic tasks.

The model is designed to allow commercial use.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.