IBM Granite 4.1 30B GGUF Quantization Released
·2026.04.30 07:37
Key point
GGUF quantization files and recommended variants for IBM Granite 4.1 30B have been released.
Details
bartowski used llama.cpp's imatrix method with a separate gist dataset to quantize IBM Granite 4.1 30B, and uploaded the resulting GGUF bundle to Hugging Face.
- It offers a wide range from BF16 57.74GB to Q8_0 30.67GB, Q6_K/Q6_K_L, Q5_K, Q4_K, IQ4, IQ3, and IQ2.
- The description marks Q6_K_L, Q6_K, Q5_K_L/M/S, Q4_K_L/M/S, IQ4_XS, Q3_K_XL, and IQ2_M as recommended variants.
- Some variants keep the embedding and output weights at Q8_0 to reinforce quality.
- It also includes a
huggingface-cli downloadexample and support notes for llama.cpp, LM Studio, koboldcpp, Jan AI, and Text Generation Web UI. - On ARM/AVX environments, online repacking of
Q4_0can be used, and the olderQ4_0_X_Xfiles are no longer needed.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.