AI Briefing
KO

GGUF NaN Explosion

·2026.04.15 05:12

Key point

The perplexity NaN cause in some MiniMax M2.7 GGUF was investigated and fixed.

Details

Investigating the perplexity NaN issue in MiniMax-M2.7 GGUF revealed that the problem wasn't unique to that particular release, appearing in 21%~38% of all GGUFs on Hugging Face.

  • GGUFs from another uploader also showed NaN in 38% (10/26), and another release showed it in 22% (5/23).
  • 99.9% KLD and other metrics were normal, so this wasn't a case of overall quality degradation.
  • Overflow in llama.cpp was pointed to as the cause.

The issue mainly occurred at block 32, with some extending to block 311. In particular, blk.61.ffn_down_exps appeared to be a key culprit, and the Q5_K and Q4_K families produced NaN at chunk 32 during PPL evaluation.

Interestingly, smaller I-quant types like IQ4_XS and IQ3_XXS had no NaN, while Q2_K_XL was fine, whereas mid-sized quantizations like Q4_K_XL caused NaN.

To summarize:

  • 10/26 NaN: bartowski release
  • 5/23 NaN: unsloth release, now fixed
  • Some quants failed at chunk 32, one at chunk 311

unsloth/MiniMax-M2.7-GGUF has now been updated to mitigate the issue, but the exact root cause of the NaN hasn't been confirmed yet, with overflow from large multiplications considered the likely cause.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.