NVIDIA Releases Kumo Tabular to Predict from Tables Without Task-Specific Training
The downloadable model uses labeled rows as context for new predictions. NVIDIA reports strong benchmark results, but warns that accuracy can slip when real tables differ from its training range.
Kumo Tabular offers teams a way to make classification and regression predictions from labeled rows without training a separate model for each dataset. NVIDIA says its synthetic-table pretraining enables this approach, but performance can weaken when real data differs from the training range or prediction rows diverge from labeled examples. The model and inference library are available now; teams still need to validate accuracy and uncertainty on held-out data, while NVIDIA’s training recipe and data generators remain forthcoming.
01
The Small, Medium, and Large versions range from 28 million to 215 million parameters; NVIDIA lists the weights on Hugging Face under its commercially usable OpenMDW 1.1 license.
02
NVIDIA reports first place on TabArena with an ELO of 1,950, but its ranking and speed results may not generalize to a particular business dataset.
03
The model handles numerical and categorical columns directly; text, images, and timestamps require preprocessing, and single-pass classification supports up to 10 classes.
A team with a labeled spreadsheet can now try a new route to predicting its next row: give the examples to Kumo Tabular instead of training a model for that particular task. NVIDIA released the open-weight model on September 29 for classification and regression, two ways to predict a category or a numerical value. The promise is a shorter path to a prediction; deciding whether that prediction is reliable remains the team’s job.
A release built around the next row
Kumo Tabular comes in Small, Medium and Large versions, ranging from 28 million to 215 million parameters. Its weights are available on Hugging Face under OpenMDW 1.1, a license NVIDIA says permits commercial use. Predictions run through NVIDIA’s open-source structured-data-models library, which downloads the weights on first use and supplies the preprocessing and other steps used in the company’s evaluations.
NVIDIA is aiming at work such as customer churn, default, demand and price prediction. Those jobs have commonly used gradient-boosted trees, with a fresh cycle of preparing data and tuning a model for each question. Kumo Tabular’s proposed shortcut is to read labeled examples at prediction time and assign labels to new rows in one pass, without training or tuning a separate model for the task. Built-in preprocessing is still part of the released workflow.
How examples become predictions
The model was pretrained entirely on artificial tables, according to NVIDIA. Its generator builds relationships among hidden variables, turns some into columns and a target to predict, then adds features meant to resemble messy data, including missing values and outliers. That training lets the released model use a new table’s labeled rows as examples rather than learn a new set of weights for each dataset.
Within a table, its Transformer considers values in their columns, how columns relate within a row, and how labeled rows relate to rows awaiting predictions. The design reuses information calculated from the labeled rows for later queries. For classification it produces class probabilities; for regression it produces quantiles that can express both a point prediction and uncertainty.
The benchmark case NVIDIA makes
NVIDIA says Kumo Tabular took the top overall position on TabArena, a leaderboard that includes tuned tree models and other tabular foundation models. Its reported speed comparison uses a uniform evaluation setup on one RTX 6000 Pro GPU. Those are company-reported benchmark results, not a guarantee that the same ranking or speed advantage will hold for a particular business dataset.
The test beyond the leaderboard
The model directly handles numerical and categorical columns. Text, images and timestamps need to be converted into features through preprocessing recipes. A single pass covers up to 10 classes; the library uses an additional method for problems with more classes. That makes the no-task-specific-training claim different from a promise that every input can be used untouched.
NVIDIA warns that accuracy may fall when a table is far outside the ranges used in training, or when the rows to be predicted differ substantially from the labeled examples. It advises checking both accuracy and how well the model’s stated uncertainty matches reality on held-out data before deployment. Skipping a training cycle, in other words, does not remove the need to test predictions against known answers.
The weights and inference library are available now, but NVIDIA says it will release the training recipe and artificial-data generators later. That leaves a concrete next step for anyone who wants to examine how its synthetic training tables were made, alongside the immediate task of testing Kumo Tabular on their own held-out data.
Sources
unite.aiNVIDIA Releases Open Kumo Tabular Model for Tabular Prediction
Reader comments
Newest comments first. Replies stay oldest first.