8B model with enhanced tool calling
Key point
IBM has released Granite-4.1-8B, enhancing tool calling and instruction-following performance.
Details
The IBM Granite Team has released Granite-4.1-8B. This model is an 8B parameter long-context instruct model, built on Granite-4.1-8B-Base and refined through SFT and reinforcement learning (RL) alignment to boost tool calling, instruction following, and chat performance.
Key information:
- License: Apache 2.0
- Supported languages: English, German, Spanish, French, Japanese, Portuguese, Arabic, Czech, Italian, Korean, Dutch, Chinese
- Primary use cases: general-purpose AI assistant, business applications, tool-use agents
- Core capabilities: summarization, classification, extraction, QA, RAG, code tasks, function calling, multilingual dialogue, FIM code completion
On benchmarks, the 8B model recorded MMLU 73.84, MMLU-Pro 55.99, BBH 80.51, IFEval 85.87, ArenaHard 68.98, HumanEval 85.37, and BFCL v3 68.27. On multilingual items, it scored MMMLU 64.84, INCLUDE 58.89, and MGSM 82.32.
The architecture is a decoder-only dense transformer, using GQA, RoPE, SwiGLU, RMSNorm, shared embeddings. The sequence length is 4096.
IBM stated that this model was trained on CoreWeave's NVIDIA GB200 NVL72 cluster, and has also provided official documentation and a GitHub repository.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.