Qwen 3.6 27B GGUF Comparison
·2026.04.28 21:18
Key point
Compared the performance of Qwen 3.6 27B's BF16, Q4_K_M, and Q8_0 quantizations.
Details
Qwen 3.6 27B was evaluated in BF16, Q4_K_M, and Q8_0 GGUF formats in a llama-cpp-python environment.
The comparison items were HumanEval (164 samples), HellaSwag (100 samples), and BFCL (400 samples), using n_ctx 32768 and checkpoint-based execution.
- BF16: average accuracy 69.78%, 15.5 tok/s, peak RAM 54GB, model size 53.8GB
- Q4_K_M: average accuracy 66.54%, 22.5 tok/s, peak RAM 28GB, model size 16.8GB
- Q8_0: average accuracy 66.15%, 18.0 tok/s, peak RAM 42GB, model size 28.6GB
In detail, Q4_K_M was presented as the most practical choice. Its BFCL score was nearly identical to BF16, and while HumanEval was about 5.5 points lower, the gap on HellaSwag was only 4 points.
To summarize:
- Q4_K_M is about 1.45x faster than BF16 and uses 48% less RAM.
- Q8_0 had a slightly higher HumanEval score than Q4_K_M, but its advantages in RAM and speed were small.
- Looking at top quality alone, BF16 was superior, while Q4_K_M was suggested as the leading option for local/CPU deployment.