NVIDIA Kumo Tabular Sets a New Accuracy-Efficiency Frontier for Tabular Prediction
Hugging Face
NVIDIA put Kumo Tabular on Hugging Face: an open model for table data that predicts without training. It claims top spots on four benchmarks, which is a loud message to old-school tree models.
Based on reporting by Hugging Face — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
NVIDIA has released Kumo Tabular, an open foundation model for tabular data that now lives on Hugging Face. Feed it a table with labeled rows and it can predict new rows in one forward pass, with no training, no tuning and no feature engineering. It handles both classification and regression, and NVIDIA says it comes in three sizes, from 28 million to 215 million parameters.
That alone is the headline. The bigger shift is the pitch underneath it: tables are supposed to work more like prompts. Kumo Tabular was pretrained only on artificial data, then taught to read context rows and answer query rows directly, the same broad idea that made large language models so useful on text. It runs through NVIDIA’s open-source structured-data-models library and is released under the OpenMDW-1.1 license for commercial use.
The model itself is a Transformer built for table structure. NVIDIA says it uses column attention, row attention and in-context attention, with special handling for missing values and separate paths for numerical and categorical data. It also uses a length-aware attention temperature so attention stays sharp as tables get bigger, plus a setup where query rows can reuse cached context calculations instead of recomputing everything each time.
Training was fully synthetic. NVIDIA built the data with Structural Causal Models, sampling endless artificial tables with different sizes, mechanisms and missingness patterns. It trained classification and regression separately, in three stages: first on tables with 1,024 rows and up to 100 columns, then with context sizes from 400 to 10,240 rows, then up to 60,000 rows. The company says the small, medium and large versions saw about 35 million, 71 million and 137 million artificial tables.
On the benchmark side, NVIDIA says Kumo Tabular ranks first on TabArena, BeyondArena, TALENT and ScoringBench. On TabArena, it posts an ELO of 1950 and runs 17 faster than LimiX-2 under a single RTX 6000 Pro evaluation setup. On BeyondArena it reaches an ELO of 1418 with an Improvability score of 7.78%, and on TALENT it tops classification accuracy, log-loss and regression RMSE. The catch is familiar: it only works on numerical and categorical columns directly, text and timestamps need preprocessing, and NVIDIA warns that accuracy can fall outside the training ranges or when query rows drift away from the context data.
My take — AI-written commentary, not fact-checked reporting
This is the part where tabular ML gets its LLM moment, and the old gradient-boosted-tree priesthood will not enjoy the sermon. Synthetic-only pretraining is a bold move, but at least NVIDIA is not pretending the model learned magic from the sky. The real test is whether teams use it because it is genuinely easier, or because “foundation model” still has a strong perfume in procurement meetings.
Read more about this at: Hugging Face
Related stories
Prior Labs Releases TabPFN-3.5: A Tabular Foundation Model That Beats the Winning Otto Kaggle Solution With Default Settings
MarkTechPost · 1 week ago ·
12
Nums AI Releases Causilo: A Tabular Foundation Model That Tops TabArena Among Single Models
MarkTechPost · 1 week ago ·
32
Introducing TabFM: A zero-shot foundation model for tabular data
Google Research · 2 months ago ·
10