TLDRocket
Sign in

IBM’s new Granite 4.2 models add reasoning and stay dense

The New Stack Frederic Lardinois Covered by 5 sources

IBM launched Granite 4.2, a new set of open-weight AI models that can reason, but stay dense and text-only. The pitch is enterprise agents that think without getting too heavy to run.

Based on reporting by The New Stack, Frederic Lardinois — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

IBM rolled out Granite 4.2 on Tuesday, and the company is still going its own way. The new family comes in 3 billion, 8 billion, and 30 billion parameter sizes, and it sticks with dense, decoder-only models that were pre-trained from scratch. That’s a different bet from the hybrid and Mamba-heavy direction a lot of the field has been taking.

This release is also a course correction of sorts. IBM tried hybrid approaches with Granite 4.0, then brought Granite 4.1 back to an all-attention, dense Transformer design, arguing that the simpler architecture was easier to fine-tune for downstream work. Now Granite 4.2 adds reasoning on top, and IBM is calling it a reasoning-focused release.

The models can switch between thinking and non-thinking modes, and there’s also a low-effort mode that uses only a small number of reasoning tokens for easier questions. That matters because IBM was skeptical of reasoning models not long ago, saying non-reasoning systems could still make sense for enterprise tasks like instruction following and tool calling. The company seems to have changed its mind, but not its taste for optionality.

These are text-only models, unlike some rivals in the same class, and IBM is pairing them with other products too. On Tuesday it also launched two new speech recognition models in the Granite Speech family. The Granite Vision 4.1 4B model already exists on the multimodal side, so a 4.2 vision update would not be a surprise, but IBM did not announce one here.

Under the hood, the family was pre-trained on 15 trillion tokens across five phases, including a long-context phase that pushes the context window to 512,000 tokens, even though the released configuration natively supports 128K. IBM also used 1 trillion tokens of synthetic code from its CodeAlchemy pipeline. The 8B and 30B models got an extra agentic reinforcement learning step for tool use, code editing, terminal work, and web search, while the 3B model supports tool calling too.

IBM is not claiming a benchmark coronation. Qwen 3.8 27B beats the Granite models across the board, especially on coding, and IBM’s own results are described as inconsistent there. But the 8B model often gets close to the 30B model, and IBM is clearly aiming at a different sweet spot: models that are light enough to run on a Mac or even some lower-end Nvidia RTX GPUs, while still being usable for high-throughput agentic work.

My take — AI-written commentary, not fact-checked reporting

IBM is making the right bet here: enterprises don’t need another oversized demo machine, they need something that can actually run all day without drama. The neat trick is not bigger reasoning, but cheaper reasoning — which is exactly the kind of boring, useful progress the AI market keeps pretending it’s above.

Read more about this at: The New Stack

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.