AI Briefing
KO

NVIDIA Releases Diffusion-Based LLM

·2026.06.25 17:34

Key point

NVIDIA has unveiled Nemotron-TwoTower-30B-A3B, a diffusion-based language model that boosts generation speed by 2.42x.

Details

NVIDIA has released Nemotron-TwoTower-30B-A3B-Base-BF16, a new language model built on the Nemotron 3 Nano 30B-A3B backbone. This model adopts a unique Diffusion-based architecture that departs from the conventional token-by-token generation approach.

Key technical features include:

  • Two-Tower Architecture: Combines a fixed Autoregressive context tower with a Diffusion Denoiser tower to fill in token blocks in parallel.
  • Performance and Efficiency: With NVIDIA's default Mask-diffusion setting, the model maintains 98.7% of benchmark quality compared to existing autoregressive models while achieving 2.42x faster generation throughput.

This model presents a new architectural approach to overcoming the generation speed limitations of existing LLMs.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.