Skip to content

Towards Universal Tabular Embeddings: A Benchmark Across Data Tasks

Apr 2026 · arXiv.org · Vol abs/2604.21696 · 4 citations · 59 references
Computer Science

TL;DR

TEmBed is introduced, the Tabular Embedding Test Bed, a unified benchmark for systematically evaluating tabular embeddings across four representation levels: cell, row, column, and table and shows that which model to use depends on the task and representation level.

Abstract

Tabular foundation models aim to learn universal representations of tabular data that transfer across tasks and domains, enabling applications such as table retrieval, semantic search and table-based prediction. Despite the growing number of such models, it remains unclear which approach works best in practice, as existing methods are often evaluated under task-specific settings that make direct comparison difficult. To address this, we introduce TEmBed, the Tabular Embedding Test Bed, a unified benchmark for systematically evaluating tabular embeddings across four representation levels: cell, row, column, and table. Evaluating a diverse set of tabular representation learning models, we show that which model to use depends on the task and representation level. Our results offer practical guidance for selecting tabular embeddings in real-world applications and lay the groundwork for developing more general-purpose tabular representation models.

View source

Similar papers

Can a single tabular embedding model service different tasks?

An initial approach that first aligns heterogeneous table-cell representations into a shared space using Hirschfeld–Gebelein–Rényi maximal correlation (HGR) is proposed, and it is found that it generalizes competitively compared to models specifically designed for individual tasks.

Unknown authors · 0 citations
#machine learning Preprint Sep 2026

Benchmarking Attention for Tabular Foundation Models

Tabular in-context learners such as TabPFN, Mitra, or ConTextTab rely on alternating row and column attention over 2D sequences of latent embeddings. These attention patterns differ markedly from the one-dimensional case in language models: row attention involves longer sequences while column attention operates on much...

Maximilian Schambach, Clemens Biehl, Sam Thelin · 0 citations
Review Open access Aug 2026

The Current Generation of Tabular Foundation Models: A Critical Review

This review is, to the authors' knowledge, the first organised around the current generation of tabular foundation models, and taxonomises the architectures by pretraining regime, maps the capability space across five axes, isolates the language-model-on-tabular strand for prediction, feature engineering and generation...

S. Kurashkin, V. Tynchenko, Alexey S. Borodulin et al. · 1 citation
#machine learning Preprint Sep 2026

TabFM-Auto: Self-Evolving Pipelines for Tabular Foundation Models

Tabular foundation models achieve strong zero-shot accuracy on structured data by pretraining on synthetic tables, but they ignore the column names, task descriptions, and auxiliary files that carry dataset semantics. Meanwhile, self-evolving machine learning engineering (MLE) agents train models from scratch on each d...

Deqing Fu, Huang-Yuan Su, Rajat Sen et al. · 1 citation
#machine learning Preprint Sep 2026

TabFM: A Zero-Shot Foundation Model for Tabular Data

Tabular machine learning typically relies on per-dataset workflows, fitting tree ensembles or running AutoML searches from scratch for every task. We present TabFM, a 400M-parameter tabular foundation model that formulates supervised tabular prediction as in-context learning. TabFM produces calibrated zero-shot predict...

Wei-Hao Kong, Erez Louidor Ilan, Shu-Xin Nie et al. · 9 citations · ⚡2
Preprint Aug 2026

Tabular foundation models for non-tabular tasks

Tabular foundation models (TFMs) have recently emerged as a promising paradigm for machine learning on tabular data, offering the ability to generalize across datasets without task-specific training. Since many machine learning datasets can be represented as tables, this raises the question: does TFM capability extend...

Goran Nakerst, John Brennan, W. Beugeling et al. · 0 citations

Related blog posts

Microsoft Research Blog Sep 30, 2026

Forecasting space weather risks on power grids

Extreme space-weather events can damage power systems on Earth and degrade GPS accuracy and satellite operations. A new machine learning system can predict where damage is likely to occur 30-60 minutes before a storm arrives. The post Forecasting space weather risks on power grids appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.