TLDRocket
Sign in

NVIDIA Releases Kumo Tabular: Open Tabular Foundation Models That Predict New Rows in a Single Forward Pass

MarkTechPost Asif Razzaq ● Covered by 2 sources

NVIDIA released Kumo Tabular, a model that predicts new table rows in one pass. The catch: it’s open enough for commercial use, and NVIDIA says it tops several benchmarks.

Based on reporting by MarkTechPost, Asif Razzaq — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

NVIDIA has a new tabular model family called Kumo Tabular, and it is built for a very specific trick: give it labeled rows as context, then ask it to predict new rows in a single forward pass. No training loop. No hyperparameter hunt. No feature wrangling first. If you’ve seen TabPFN or TabICL, the basic idea will feel familiar, but NVIDIA is pushing it through its own structured-data-models library and giving it a commercial-friendly license.

The family comes in Small, Medium, and Large versions, with roughly 28M to 215M parameters. NVIDIA says the weights are under OpenMDW-1.1, while the SDM library itself is Apache-2.0. The setup expects Python 3.11+ and PyTorch 2.7+, and the examples are aimed at a CUDA GPU. So this is not a research-only toy sitting on a shelf somewhere. It is meant to run.

Under the hood, Kumo Tabular is a Transformer tuned for tables. It uses attention at the column, row, and in-context levels, with separate handling for numerical and categorical values, missing values, and large tables. The model also outputs 999 quantiles for regression, which gives a point estimate and an uncertainty signal instead of just a single number. NVIDIA also says the attention stays sharper as tables get larger by scaling queries with a learned temperature tied to key count.

The training story is synthetic all the way down. Kumo Tabular was pretrained on artificial tables generated from structural causal models, with random causal graphs and messy details like missing values, high-cardinality categories, heavy-tailed targets, and duplicate rows. NVIDIA trained it in three stages, with context growing from 1,024 rows to 60,000 rows and up to 100 columns. The company says classification and regression are separate models, and it plans to release the training recipe and data generators soon.

On benchmarks, NVIDIA is making a loud claim. With default settings, it says Kumo Tabular ranks first overall on TabArena with an Elo of 1950, and it also leads BeyondArena, TALENT, and ScoringBench in the reported results. The bigger practical punchline may be the license, though: Kumo Tabular and TabICLv2 are the permissive options here, while TabPFN-3, LimiX-2, and TabFM carry non-commercial terms.

My take — AI-written commentary, not fact-checked reporting

Open weights still matter, and this is why. Benchmarks are nice, but a commercial-friendly license does more to move real projects than another leaderboard trophy ever will. The industry keeps acting shocked when permission beats performance marketing.

Read more about this at: MarkTechPost

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.