SyncAI.news, a Varaisys broadcasting
NVIDIA Kumo Tabular Sets a New Accuracy-Efficiency Frontier for Tabular Prediction
HF

Hugging Face Blog

· 1 min read

AI LabsHugging Face Blog

NVIDIA Kumo Tabular Sets a New Accuracy-Efficiency Frontier for Tabular Prediction

Highlights (TL;DR)

NVIDIA Kumo Tabular, part of the NVIDIA Kumo Structured model collection, is an open foundation model for tabular data now available on Hugging Face. Given a table of labeled rows, it predicts the labels of new rows in a single forward pass, with no training, no tuning, and no feature engineering, for both classification and regression. It was pretrained only on artificial data, comes in three sizes (28M to 215M parameters), runs through our open-source library, and is released under the OpenMDW-1.1 license for commercial use. It ranks first on the four benchmarks TabArena, BeyondArena, TALENT and ScoringBench.

  • Model Code: https://github.com/NVIDIA/structured-data-models
  • Model Weights: https://huggingface.co/nvidia/Kumo-Tabular

The Shift to Tabular Foundation Models

Tabular data is the backbone of enterprise machine learning. Customer records, transactions, sensor logs, claims, and orders all live in tables, and predicting churn, default, demand, or price from them is among the most common machine learning tasks in industry. For two decades, this work has been done with gradient-boosted trees, and it has worked well. But the lifecycle around those models has barely changed. Every new question means collecting labels, engineering features, searching hyperparameters, validating, and deploying a model that knows nothing about tables in general and learns each task from scratch.

Large Language Models showed a different way of working with new tasks. Given a few examples in the prompt, a pretrained model solves the task without updating a single weight. This is in-context learning, and it applies to tables just as well as to text: a model pretrained on millions of tables can read a labeled table as its context and predict the labels of new rows directly.

How Kumo Tabular Works

How Kumo Tabular was Built

Kumo Tabular is pretrained entirely on artificial tables. Each training table is sampled from a Structural Causal Model (SCM) in the six steps shown below:

Original source

This story was published by Hugging Face Blog. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on huggingface.co

Similar News