AI Briefing
KO

llama.cpp merges SM120 NVFP4 MMQ

·2026.04.29 09:39

Key point

llama.cpp has merged initial native NVFP4 MMQ support for SM120, and related GGUF conversions have also been released.

Details

A native NVFP4 MMQ path for SM120 has been merged into llama.cpp's CUDA backend. The PR was bundled as 23 commits and merged on April 28.

The changes can be summarized as adding NVFP4 kernels for Blackwell, unifying the FP4 quantization path, and cleaning up the FP4 MMA templates.

Related GGUF conversions were also shared right away.

  • NVFP4 GGUF for Gemma 4 31B IT
  • NVFP4 GGUF for Nemotron Cascade 2 30B A3B
  • NVFP4 GGUF for Qwen3.5 35B A3B

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.