AI Briefing
KO

llama.cpp Adds Support for NVIDIA Nemotron-3-Puzzle-75B-A9B Model

·2026.09.03 16:35

Key point

llama.cpp now supports running the Nemotron-3-Puzzle-75B-A9B model, which adopts NVIDIA's hybrid MoE architecture.

Details

Support for NVIDIA's Nemotron-3-Puzzle-75B-A9B model has been added to the llama.cpp project. This model utilizes a 75B MoE architecture and is currently runnable, though MTP (Multi-Token Prediction) functionality is not yet supported.

Model Architecture and Features

  • Hybrid Structure: Uses a hybrid MoE architecture with alternating Mamba, MoE, and Attention layers.
  • MTP Support Plan: Designed to support MTP functionality to increase text generation speed, similar to the parent model Nemotron-3-Super, but it has not yet been implemented in llama.cpp.
  • Parameter Efficiency: Reduced total parameters from 120.7B to 75.3B and active parameters from 12.8B to 9.3B compared to the parent model, achieving a lighter footprint.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.