hipfire boosts HFQ4 prefill by 3x
Key point
Adding opt-in MMQ to HFQ4 prefill made Strix Halo speed more than 3x faster.
Details
hipfire has added an opt-in MMQ path for HFQ4-G256 prefill.
It's enabled via HIPFIRE_MMQ=1, and targets gfx1100, gfx1101, gfx1102, gfx1103, gfx1150, gfx1151.
The implementation pre-quantizes prefill activations into a Q8_1 MMQ layout, and uses i8 WMMA with 128×128 output/batch tiles and LDS staging.
On Strix Halo / gfx1151 running Qwen3.5 9B HFQ4/MQ4, long-prefill throughput rose from the existing 310~340 tok/s to around 1140~1260 tok/s.
q81024pp: 328.9 → 1222.7 tok/s (3.72x)asym31024pp: 329.9 → 1259.1 tok/s (3.82x)- Across the full table, speedups of about 3.0~3.9x were confirmed depending on KV mode.
It's not the default, and it's similar in form to llama.cpp's AMD MMQ prompt-processing path.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.