Quantization Ranking Battle
·2026.04.16 18:38
Key point
Compared MMLU scores of GGUF quantized models in a 24GB VRAM environment.
Details
Compared GGUF quantized models in a llama.cpp environment based on MMLU subset (DEV+TEST).
- Settings: ctx 8192, seed 42, fa on
- The top tier showed almost no difference between Qwen3.5-27B-UD-Q5_K_XL.gguf 87.33% and Qwen3.5-27B-UD-Q4_K_XL.gguf 87.25%.
- Next was Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled.i1-Q4_K_M.gguf 87.02%.
- Qwen3-Coder-Next-UD-Q4_K_XL.gguf followed at 84.38%, and Qwen3.5-35B-A3B-Uncensored-HauhauCS-Aggressive-Q4_K_M.gguf recorded 83.25%.
- The smaller Qwen3.5-9B-UD-Q8_K_XL.gguf scored 78.81%, and gemma-4-31B-it-UD-Q4_K_XL.gguf scored 78.36%, with errors=1 noted.
- The largest model, Qwen3.5-397B-A17B-UD-IQ2_XXS-00001-of-00004.gguf, was relatively low at 65.80%.
This is a benchmark comparison that shows at a glance the MMLU performance differences across various GGUF quantization combinations under the same conditions.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.