AI Briefing
KO

Gemma4 tops local benchmark

·2026.05.01 04:59

Key point

In an 18GB M3 Pro bench, Gemma4:e4b ranked first with 72%, while Qwen3.5:9b scored 68%.

Details

Bench 3 was run again on an 18GB M3 Pro. Reflecting the limitations of Bench 2, new prompts and harder tasks were added, and all models were given a 4096 token budget. Supported models had think:false applied, and sparse models had their active-params shown separately. Evaluation covered 25 tasks (10 finance, 7 reasoning, 8 code), run 3 times per model and aggregated by median, with temp=0, seed=42, and max_tokens=4096 fixed. tokens-per-correct was also tracked as a key metric.

  • Lineup: gemma4:e4b 9.6GB (4B active MoE), qwen3.5:9b 6.6GB, granite4:3b 2.1GB, olmo-3:7b-think-q8_0 7.8GB, nemotron-3-nano:4b 2.8GB
  • Grading: Same local machine, same deterministic grader, no LLM-as-judge

gemma4:e4b was the best overall at 72%. It showed the most even general-purpose performance with 70% in finance, 71% in reasoning, and 75% in code.

qwen3.5:9b was strongest in finance at 80%, and also scored 75% in code. It ranked second overall at 68%, and the 1024 token cap issue pointed out in Bench 2 was resolved this time with the 4096 token budget, normalizing its performance.

Next were granite4:3b at 44%, olmo-3:7b-think-q8_0 at 40%, and nemotron-3-nano:4b at 36%. These results show not just how the models compare against each other, but also how much output budget and reasoning settings can change how a benchmark is interpreted.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.