IBM Releases Granite 4.2 Reasoning LLMs: 3B–30B, 512K Context, and Agentic RL
Key point
IBM released three Granite 4.2 reasoning LLMs featuring 512K context and agentic RL.
Details
IBM has released the Granite 4.2 LLM series. This version consists of dense (decoder-only) reasoning models in three sizes: 3B, 8B, and 30B, with all models distributed under the Apache 2.0 license.
The models were pre-trained from scratch on approximately 15T tokens following a five-stage strategy. Through this process, the context window was expanded to 512K tokens. Subsequently, they underwent supervised fine-tuning (SFT) with chain-of-thought (CoT), reasoning, and agent trajectory data.
Training Pipeline and Features
In the post-training stage, a multi-stage reinforcement learning (RL) pipeline was applied. Notably, the 8B and 30B models were trained on agentic RL, learning to act using tools in real sandbox environments. All models can switch between thinking and non-thinking modes based on task difficulty, and support a low-effort reasoning mode that uses a short reasoning budget for easy questions. Additionally, native tool calling capabilities are built in.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.