AI Briefing
KO

DFlash Released for Gemma 4

·2026.05.08 23:18

Key point

z-lab has released the DFlash drafter model and paper for Gemma 4.

Details

z-lab has released gemma-4-26B-A4B-it-DFlash. This model is a block diffusion-based speculative decoding drafter, and it must be used together with google/gemma-4-26B-A4B-it.

  • Released materials: arXiv paper, GitHub repository, project blog, Hugging Face model card
  • Serving support: provides integration examples for vLLM and SGLang
  • Performance claim: presents up to 3.7x throughput improvement on a single B300 GPU environment
  • Measurement conditions: includes benchmarks based on thinking enabled, greedy decoding, max output length of 4096, and concurrency of 1/8/32

In the benchmarks, it showed higher generated tokens/sec compared to AR across Math500, GSM8K, HumanEval, MBPP, and MT-Bench, and acceptance length was also presented per task. In vLLM, Gemma 4 DFlash support has not yet been merged, so installation via PR is required, and SGLang can also be run through an experimental option.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.