Q4_K_XL faster than Q4_K_M
·2026.04.26 18:51
Key point
In an 8GB VRAM / 32GB RAM environment, Q4_K_XL was slightly faster than Q4_K_M and also produced fewer output tokens.
Details
In a benchmark running Qwen3.6 35b a3b tuned for an 8GB VRAM / 32GB RAM environment, Unsloth's Q4_K_XL showed slightly better performance than Q4_K_M.
- Settings: CtxSize 131,072, GpuLayers 99, CpuMoeLayers 38, Threads 16, BatchSize/UBatchSize 4096/4096, K/V cache q8_0
- Comparison results:
- Avg Tokens/sec: 28.92 → 29.78
- Median Tokens/sec: 30.96 → 32.08
- Avg Wall Seconds: 108.03s → 99.93s
- Avg Output Tokens: 3,031.8 → 2,895.8
- Input processing and decoding speed also improved slightly, with input token processing speed increasing from 50.20 → 55.96 tok/s.
- The author explained that since the first run includes initialization time for loading the model from storage into RAM, measurements were taken 5 times based on actual usage.
The key point is that Q4_K_XL, which uses more memory, was about 3% faster than Q4_K_M and also produced about 4.5% fewer output tokens in this environment.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.