Gemma 4 26B NVFP4 Released
·2026.05.01 12:37
Key point
NVIDIA released an NVFP4 quantized checkpoint of the Gemma 4 26B model.
Details
NVIDIA uploaded an NVFP4 quantized checkpoint based on Gemma 4 26B A4B IT to Hugging Face.
- The base model is Google DeepMind's Gemma 4 26B IT, which supports text and image input.
- The model card highlights 256K context and support for 140+ languages.
- This version was quantized with NVIDIA Model Optimizer and is provided for vLLM inference.
- The documentation lists the model size as 14B params, with total parameters of 25.2B and active parameters of 3.8B.
- Evaluations presented figures including GPQA Diamond 79.9%, AIME 2025 90.0%, MMLU Pro 84.8%, and LiveCodeBench 79.8%.
- However, it currently only supports TP=1 in vLLM, and constraints related to EP and Flashinfer-TRTLLM along with unresolved issues are also noted.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.