TLDRocket
Sign in

IBM releases SOTA Granite Time Series PatchTST-FM-r2 model with commercial-friendly license

Hugging Face

IBM just released Granite Time Series PatchTST-FM-r2, a zero-shot forecasting model with open commercial licensing. It’s the top permissively licensed zero-shot model on GIFT-Eval, so businesses can use it without the usual lock-in.

Based on reporting by Hugging Face — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

IBM has released Granite Time Series PatchTST-FM-r2, the newest model in its Granite TSFM line, and the pitch is simple: better forecasting without having to train a custom model for every dataset. The model is built for zero-shot use, which means it can generate forecasts from new series straight away, with no fine-tuning required.

The headline number is about 385 million parameters, but the more interesting part is the mix of features packed into that size. PatchTST-FM-r2 supports context lengths up to 8,192 steps, can produce flexible forecast lengths, and adds probabilistic output through a 99-quantile prediction head. It also handles missing values, which matters in real-world time series more than benchmark charts tend to admit.

On GIFT-Eval, IBM says the model ranks second among replicable zero-shot models for both CRPS and MASE as of September 8, 2026. Among permissively licensed models, it leads that category. The model is dual licensed under Apache 2.0 and OpenMDW 1.0, and IBM says the weights, architecture, inference pipeline, and reproduction code are all available.

The architecture is not a simple tweak of the earlier PatchTST-FM-r1. IBM replaced standard transformer blocks with conformer-style blocks that mix self-attention and temporal convolution, added overlapping patches with Hamming-window weighting, expanded the backbone from 20 to 30 blocks, and introduced normalization for stability. The result is meant to catch both short and long relationships in a time series, instead of forcing attention to do all the work.

IBM also lays out the training corpus, which includes selected GiftEvalPretrain datasets, synthetic data based on KernelSynth, a TSMixup corpus built from datasets outside the GIFT-Eval evaluation set, and about 500,000 synthetic CauKer sequences of length 4,096. That transparency is part of the story here. For enterprise users, a permissive license is nice, but knowing what went into the model is often the real selling point.

My take — AI-written commentary, not fact-checked reporting

This is the kind of release that makes open models matter more than the marketing noise around them. IBM is not just chasing leaderboard vanity; it’s pairing solid performance with a license companies can actually live with, which is annoyingly rare and therefore useful. The bigger lesson is that “open” starts to mean something only when the weights, code, and training story are all on the table.

Read more about this at: Hugging Face

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.