q1_0 speed explosion
·2026.04.21 20:41
Key point
By optimizing the q1_0/q8_0 dot path, up to 67x throughput improvement was reported on x86.
Details
Reworked the CPU q1_0 dot product path in llama.cpp.
- The generic implementation handles
(q1_0, q8_0)dot more efficiently by reducing bit operations and multiplications. - The x86 SIMD implementation broadly covers realistic x86_64 targets from SSSE3 through AVX2.
- Verification was carried out with
test-quantization-fns, model behavior, wikitext-2 perplexity, andllama-bench.
On Bonsai 1.7B, performance improved significantly on an AMD Ryzen 5 7640HS (65W), WSL VM, LPDDR5 6400MT, 10 threads.
- pp 512: initial 2.05 t/s → AVX2 + FMA 131.03 t/s
- tg 128: initial 1.32 t/s → AVX2 + FMA 73.85 t/s
- Progressive improvements were confirmed in the order SSSE3, AVX, AVX + F16C, AVX2 + FMA, AVX512.
- In terms of accuracy, top token match rate and KLD remained stable overall, with no significant quality degradation observed.
The author also mentioned future possibilities including nrc==2 branching, AVX512 for Zen 5/modern Xeon, and RISC-V SIMD extensions.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.