Granite 4.2 LLMs: How They're Built
IBM's Granite Team published a technical walkthrough of how the Granite 4.2 reasoning model family was built. Granite 4.2 is described as the team's first family of dense, decoder-only reasoning LLMs, released in three sizes: 3B, 8B, and 30B. All models are released under the Apache 2.0 license.
The models are post-trained from Granite-4.1 base models, which were pre-trained from scratch on roughly 15 trillion tokens using a five-phase strategy that extends the context window to 512K tokens.
Granite 4.2 adds explicit reasoning: every model can produce a chain of thought before its answer and run in thinking or non-thinking mode, with a low-effort mode between the two. All support native tool calling and can be served through an OpenAI-compatible endpoint.
Supervised fine-tuning used a mixture of agentic (31.6%) and non-agentic (68.4%) data, totaling about 7.2 million samples or roughly 100B tokens, with about 65B trainable. Quality control included normalization into OpenAI Chat format, LLM-based judges (GPT-OSS-120B and Gemma 4), heuristic rules, and SHA-256 based local and global deduplication. The 30B model had a second SFT phase for agentic coding at a learning rate of 3.0e-6.
After SFT, a multi-stage RL pipeline runs per stage with asynchronous GRPO, warm-starting from the previous checkpoint: RLVR, skill boosters, SWE agent, terminal, search, and RLHF. Agentic RL stages use real environments with sparse outcome rewards and run only for the 8B and 30B models; the 3B takes a shortened path without the agentic block. RL infrastructure pairs NeMo-RL for training with NeMo-Gym for rollouts. The post states Granite 4.2 was evaluated across agentic coding, agentic and tool use, reasoning, chat and instruction following, and long context, with supported languages including English, German, Spanish, French, Japanese, Portuguese, Arabic, Czech, Italian, Korean, Dutch, and Chinese.
Based on reporting from the original publisher. Visit the source for full context and later updates.