Qwen 3.6 KV Cache Benchmark
·2026.04.29 02:17
Key point
Measured Qwen 3.6 KV cache performance differences by format from 0 to 1M context on M5 Max.
Details
On an M5 Max 128GB, KV cache for Qwen 3.6-35B-A3B was compared using llama-bench to measure performance curves for f16, q8_0, turbo3, and turbo4 across 0 to 1M context.
- Ran with a symmetric setup matching K and V to the same type (
-ctk,-ctv), using thefeature/turboquant-kv-cachebranch combined withGGML_METAL=ON. - Each cell was repeated 3 times, and the sweep took about 8 hours with flash-attn and
mlockenabled. - At 0 depth, f16 was fastest with prefill 2962 tok/s and decode 89.4 tok/s, while turbo3's decode was about 10% slower at 79.5 tok/s.
- At 128K, turbo3 prefill reached 253 tok/s, surpassing q8_0's 245 tok/s. As context grew longer, smaller cache eased the bandwidth burden.
- The OOM limit varied significantly by cache format. f16 hit OOM at 256K, q8_0 hit OOM at 512K, while turbo3 worked up to 1M, recording prefill 30 tok/s and decode 6.5 tok/s.
- turbo3 and turbo4 had different bottlenecks. At 256K prefill, turbo3 was faster at 128 tok/s compared to turbo4's 101 tok/s, but for decode at 512K, turbo4 (16.0 tok/s) outpaced turbo3 (13.3 tok/s).
As context length grew extremely long, KV cache compression simultaneously governed both performance and memory limits, and the optimal cache type differed depending on whether it was prefill or decode.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.