Qwen3.6 larger quants are faster
·2026.04.25 06:49
Key point
On a 3070 8GB setup, Qwen3.6-35B ran faster with larger quants.
Details
Running Qwen3.6-35B-A3B-UD on an RTX 3070 8GB and DDR4 64GB setup, larger quants performed better than expected.
- IQ4_XS.gguf, at about 18GB, recorded 25~30 tok/s at 32k context.
- The larger Q4_K_XL.gguf, at about 23GB, was faster at 32 tok/s with 128k context.
- In the end, Q5_K_S offered the best balance of quality and speed, delivering about 30 tok/s.
- Even at 50k context, speed stayed above 25 tok/s.
The author concluded that for MoE models, even when VRAM looks tight, a higher quant than expected can be the better choice.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.