AI Briefing
KO

Gemma 4 MTP Released

·2026.05.06 01:01

Key point

An MTP draft model for Gemma 4 has been released, boosting inference speed by up to 2x.

Details

MTP (Multi-Token Prediction) draft model checkpoints for the Gemma 4 family have been uploaded to Hugging Face.

  • The released models are 31B-it-assistant, 26B-A4B-it-assistant, E4B-it-assistant, and E2B-it-assistant.
  • MTP is a speculative decoding approach in which a smaller, faster draft model predicts multiple tokens first, and the target model verifies them in parallel.
  • According to the model card, this approach delivers up to 2x faster decoding while maintaining the same quality as standard generation.
  • These checkpoints are aimed at low-latency inference and on-device use.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.