Gemma-4 combo, 200tok/s
Key point
On an RTX 5090, a Gemma-4 combo recorded 120~200 tok/s.
Details
In a test combining Gemma-4-31B-it Q6_K_L and Gemma-4-E2B-it Q8_0 via speculative decoding on an RTX 5090, output speed came out at around 130~200 tok/s.
- The intended use was non-agentic LLM tasks such as legal document citation extraction, classification, and title conversion.
- Input context was typically 2K~6K tokens, ranging up to 8K tokens at most.
- The setup used 31.5GB VRAM, nearly filling up the GPU.
- The author stated that for the same tasks, it produced better quality and was faster than Gemini 2.5 Flash-lite, and could be run locally.
While based on early small-scale testing, they assessed that for lightweight workflows where non-English language processing and structured JSON responses matter, this could serve as a substitute without needing a cloud API.
The accompanying example runs llama-server to load bartowski/google_gemma-4-31B-it-GGUF:Q6_K_L.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.