AI Briefing
KO

Qwen 3.6 27B GGUF Comparison

·2026.04.28 21:18

Key point

Compared the performance of Qwen 3.6 27B's BF16, Q4_K_M, and Q8_0 quantizations.

Details

Qwen 3.6 27B was evaluated in BF16, Q4_K_M, and Q8_0 GGUF formats in a llama-cpp-python environment.

The comparison items were HumanEval (164 samples), HellaSwag (100 samples), and BFCL (400 samples), using n_ctx 32768 and checkpoint-based execution.

  • BF16: average accuracy 69.78%, 15.5 tok/s, peak RAM 54GB, model size 53.8GB
  • Q4_K_M: average accuracy 66.54%, 22.5 tok/s, peak RAM 28GB, model size 16.8GB
  • Q8_0: average accuracy 66.15%, 18.0 tok/s, peak RAM 42GB, model size 28.6GB

In detail, Q4_K_M was presented as the most practical choice. Its BFCL score was nearly identical to BF16, and while HumanEval was about 5.5 points lower, the gap on HellaSwag was only 4 points.

To summarize:

  • Q4_K_M is about 1.45x faster than BF16 and uses 48% less RAM.
  • Q8_0 had a slightly higher HumanEval score than Q4_K_M, but its advantages in RAM and speed were small.
  • Looking at top quality alone, BF16 was superior, while Q4_K_M was suggested as the leading option for local/CPU deployment.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.