AI Briefing
KO

oQ vs Q vs MXFP vs UD KLD Comparison

·2026.04.30 05:37

Key point

On Qwen3.6-35B-A3B, KLD and memory usage were compared across oQ, Q, MXFP, and UD MLX.

Details

Measurements were taken on Qwen3.6-35B-A3B using mlx-vlm v0.4.4.

  • Setup: A single prompt of 16,384 tokens was fed into the bfloat16 model, using Aes Sedai's combined_all_micro.txt.
  • oQ: oQ2 was 0.277344 nats / 11.40 GiB, oQ4 was 0.028076 / 18.83 GiB, oQ6 was 0.008057 / 26.51 GiB, and oQ8 was 0.005219 / 34.27 GiB.
  • Q·MXFP: Q2 collapsed significantly at 3.093750 / 10.10 GiB, while Q4 was 0.062500 / 18.17 GiB, Q6 was 0.009094 / 26.23 GiB, and Q8 was 0.005402 / 34.30 GiB. MXFP4 was 0.111328 / 17.16 GiB, and MXFP8 was 0.041992 / 33.29 GiB.
  • UD MLX: 3-bit was 0.048584 / 15.35 GiB, and 4-bit was 0.016357 / 19.32 GiB.

Numerically, oQ4 shows lower KLD than both Q4 and MXFP4, and oQ6/oQ8 come very close to Q6/Q8.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.