Qwen 3.6 Wins Big
Key point
In a 30k LOC agentic evaluation, Qwen 3.6 35B outperformed Gemma 4 26B.
Details
In a personal evaluation harness, Qwen3.6-35B-A3B and Gemma 4-26B-A4B were compared under the same conditions.
The harness has the model agentically fix 37 intentionally introduced issues in a roughly 30k LOC repository, and also includes a task to extract, summarize, and evaluate key information from a 40-60 page PDF.
The evaluation covered the following five categories.
- Agentic capabilities
- Coding
- Image-to-text synthesis
- Instruction following
- Reasoning
Both were run with UD-Q4_K_XL quantization and optimal sampling parameters, and Gemma 4 had the latest chat template fixes applied along with the -cram and -ctkcp flags.
The results favored Qwen3.6.
- Tests fixed: 32/37 (86.5%)
- Regressions: 0
- Net score: 32
- Post-run failures: 5
- Duration: 49 minutes
Gemma 4 results were
- Tests fixed: 28/37 (75.7%)
- Regressions: 8
- Net score: 20
- Post-run failures: 17
- Duration: 85 minutes
The author noted that Qwen judged the remaining 5 failures to be out of scope and skipped them, while Gemma gave up partway through.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.