AI Briefing
KO

Qwen KV Cache Benchmark

·2026.04.30 05:00

Key point

Re-measured the quality and long-context speed of Qwen 3.6-35B-A3B KV cache quantization on M5 Max.

Details

Overnight, re-measured KV cache quantization quality and speed on M5 Max under the same Qwen 3.6-35B-A3B Q8 and the same TheTom TurboQuant fork (feature/turboquant-kv-cache) conditions.

Quality evaluation (wikitext-2, context 4096)

  • At 512 context, KV cache differences weren't clearly visible, so the comparison was done at context 4096 to see the effect.
  • After saving the f16 baseline with --kl-divergence-base, compared PPL, KL divergence, and top-1 token agreement for each quantized cache.
  • q8_0 recorded PPL 5.7433, KL 0.0016, top-1 98.64%, essentially no difference from f16 (PPL 5.7438).
  • turbo3 recorded PPL 5.8092, KL 0.0199, top-1 93.93%, and turbo4 recorded PPL 5.7810, KL 0.0131, top-1 95.28%.
  • As the compression ratio increased, quality loss grew as well, but q8_0 KV was nearly lossless at this depth.

Asymmetric K/V (depth sweep)

  • Based on decode tok/s, the q8_0 K / turbo4 V combination performed best: at 0 / 8K / 32K / 128K / 256K / 512K, it recorded 82.9 / 75.4 / 66.0 / 41.0 / 27.1 / 16.5 tok/s respectively.
  • q8_0 K / turbo3 V was slightly slower than that, and f16 K / turbo4 V slowed down significantly at deeper contexts, so 256K and above were omitted.
  • Notably, q8_0 K / turbo4 V approached the throughput level of the previous symmetric q8_0 at 256K, and it worked all the way to 512K, where symmetric q8_0 had OOM'd, showing standout memory efficiency in long-context scenarios.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.