Unsloth Dominates
Key point
On Gemma 4 26B-A4B GGUF, Unsloth quants top the charts by KLD in most cases.
Details
By Mean KLD, nearly all Unsloth GGUF models landed on the Pareto frontier, achieving top performance in 21 of 22 sizes.
The core metric shows how well a quantized model preserves the BF16 output distribution, and this result showed a similar trend in other metrics such as 99.9% KLD.
Additionally, the Q6_K quant was updated with a more dynamic approach, and there's no need to re-download the existing version. The same kind of improvement was also applied to Qwen3.6.
A new UD-IQ4_NL_XL quant was also released.
- At 14.6GB, it's suitable for 16GB VRAM
- Positioned between UD-IQ4_XS(13.4GB) and UD-Q4_K_S(16.4GB)
The MLX quant was also adjusted to be more dynamic. In the published comparison table, the new version showed a slight improvement over the previous version.
- Perplexity: 4.772 → 4.766
- Mean KLD: 0.0177 → 0.0163
- 99.9% KLD: 0.8901 → 0.8398
- Disk Size: 21.4GB → 21.6GB
Detailed graphs and GGUF files are available in the Gemma 4 and Qwen3.6 benchmark documents and Hugging Face repository.