AI Briefing
KO

Qwen 3.6 Wins Big

·2026.04.18 05:15

Key point

In a 30k LOC agentic evaluation, Qwen 3.6 35B outperformed Gemma 4 26B.

Details

In a personal evaluation harness, Qwen3.6-35B-A3B and Gemma 4-26B-A4B were compared under the same conditions.

The harness has the model agentically fix 37 intentionally introduced issues in a roughly 30k LOC repository, and also includes a task to extract, summarize, and evaluate key information from a 40-60 page PDF.

The evaluation covered the following five categories.

  • Agentic capabilities
  • Coding
  • Image-to-text synthesis
  • Instruction following
  • Reasoning

Both were run with UD-Q4_K_XL quantization and optimal sampling parameters, and Gemma 4 had the latest chat template fixes applied along with the -cram and -ctkcp flags.

The results favored Qwen3.6.

  • Tests fixed: 32/37 (86.5%)
  • Regressions: 0
  • Net score: 32
  • Post-run failures: 5
  • Duration: 49 minutes

Gemma 4 results were

  • Tests fixed: 28/37 (75.7%)
  • Regressions: 8
  • Net score: 20
  • Post-run failures: 17
  • Duration: 85 minutes

The author noted that Qwen judged the remaining 5 failures to be out of scope and skipped them, while Gemma gave up partway through.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.