AI Briefing
KO

IBM Unveils Granite 4.1 8B

·2026.05.01 09:28

Key point

IBM released Granite 4.1 3B, 8B, and 30B, with the 8B rivaling 32B-class models.

1 / 2

Details

IBM unveiled Granite 4.1 3B, 8B, and 30B. All three models use a dense decoder-only architecture and the Apache 2.0 license, are designed for enterprise use, and were trained on 15 trillion tokens.

The key model is 8B. Despite being much smaller than the previous generation's Granite 4.0-H-Small (32B MoE, 9B active), it scored higher on several benchmarks.

  • ArenaHard 69.0
  • BFCL V3 68.3
  • GSM8K 92.5
  • EvalPlus 80.2

The 3B also posted solid numbers for its size — IFEval 82.1, GSM8K 87.0, BFCL V3 60.8 — while the 30B ranked at the top of IBM's chart with BFCL V3 73.7, ArenaHard 71.0, and GSM8K 94.2.

The training pipeline was refined in stages. Pretraining proceeded through 5 stages, with the proportion of code and math significantly increased from the midpoint onward. Before SFT, low-quality responses were filtered out using LLM-as-Judge and rule-based filters, leaving only 4.1M samples. RL was then run in 4 stages; when general-chat RLHF degraded math performance, a separate math RL stage was used to recover it.

Long context was extended in the order 32K → 128K → 512K. The 8B and 30B ultimately support up to 512K, while the 3B only goes up to 128K. On RULER, the 8B scores declined as 83.6→79.1→73.0 and the 30B as 85.2→84.6→76.7 — scores decrease gradually as context lengthens, but there's no sharp drop-off.

For deployment, the models support Ollama, vLLM, Transformers, and the IBM API, and FP8 quantization is also available. However, since the comparison figures come from IBM's own evaluation harness, cross-verification is needed. Even so, this stands out as a notable case of a smaller dense model outperforming a larger MoE model across multiple tasks.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.