On September 29, NVIDIA published Kumo Tabular on Hugging Face, part of the Kumo Structured collection. Given a table of labeled rows, it predicts the labels of new rows in one forward pass, for classification or regression, with no further training, tuning, or feature engineering. It was pretrained only on artificial data, in three sizes from 28 million to 215 million parameters. It runs through an open-source library and is released under the OpenMDW-1.1 license for commercial use. Code is at NVIDIA/structured-data-models. Weights are at nvidia/Kumo-Tabular.

[1]
Chalk on slate: a grid of empty cells with one ochre mark in the lower row.
The grid is empty except for one marked cell. That is a new row labeled from the rows already in the table. An illustration, not a data sheet., AI-generated illustration, not a news photograph

It is a Transformer built around a table, using column attention, row attention, and in-context attention. A prediction has to see what a value means inside its column, how the columns of a row interact, and how labeled context rows relate to unlabeled query rows. Numbers and categories use different weights. Missing values are not imputed. Query rows attend only to the context, not to other queries, so the context keys and values can be computed once and reused. Classification returns class probabilities. Regression returns 999 quantiles, and from those a point prediction and an uncertainty estimate.

Every training table is an artificial sample from a structural causal model. Classification and regression are separate models. The small, medium, and large sizes saw about 35, 71, and 137 million artificial tables. The training recipe and the generators are described as coming soon. They were not released with the post.

[1]

The post says all three sizes were run at default settings against the TabArena board, including tuned gradient-boosted trees, AutoGluon, and newer tabular foundation models. Kumo Tabular ranks first overall, with an ELO of 1950 on TabArena. On BeyondArena the ELO is 1418 and the Improvability score is 7.78%, also first. On TALENT the average ranks for classification accuracy, classification log-loss, and regression RMSE are 6.67, 3.98, and 4.22, and the overall rank is first. On ScoringBench, Large and Medium are first and second by average rank. Those ranks are the post’s account. This piece did not open the boards.

The limits are stated. It handles numerical and categorical columns directly. Text, images, or timestamps have to be turned into features by the library’s recipes. One forward pass covers up to 10 classes. More classes use error-correcting output codes. Accuracy may fall on tables far outside the training range, or when query rows come from a different distribution than the context. Accuracy and calibration should be checked on held-out data before deployment.

[1]

要点

  • Labeled rows are the context. One forward pass classifies or regresses new rows, with no extra training or feature engineering.
  • Pretraining used only artificial tables. Three sizes run from 28 million to 215 million parameters, under OpenMDW-1.1.
  • The post says it ranks first on TabArena, BeyondArena, TALENT, and ScoringBench. The TabArena ELO is 1950.
  • Only numbers and categories are handled directly. One pass covers up to 10 classes. A shift in distribution has to be checked.