oQ vs Q vs MXFP vs UD KLD Comparison
·2026.04.30 05:37
Key point
On Qwen3.6-35B-A3B, KLD and memory usage were compared across oQ, Q, MXFP, and UD MLX.
Details
Measurements were taken on Qwen3.6-35B-A3B using mlx-vlm v0.4.4.
- Setup: A single prompt of 16,384 tokens was fed into the
bfloat16model, using Aes Sedai'scombined_all_micro.txt. - oQ:
oQ2was 0.277344 nats / 11.40 GiB,oQ4was 0.028076 / 18.83 GiB,oQ6was 0.008057 / 26.51 GiB, andoQ8was 0.005219 / 34.27 GiB. - Q·MXFP:
Q2collapsed significantly at 3.093750 / 10.10 GiB, whileQ4was 0.062500 / 18.17 GiB,Q6was 0.009094 / 26.23 GiB, andQ8was 0.005402 / 34.30 GiB.MXFP4was 0.111328 / 17.16 GiB, andMXFP8was 0.041992 / 33.29 GiB. - UD MLX:
3-bitwas 0.048584 / 15.35 GiB, and4-bitwas 0.016357 / 19.32 GiB.
Numerically, oQ4 shows lower KLD than both Q4 and MXFP4, and oQ6/oQ8 come very close to Q6/Q8.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.