A systematic evaluation of table-level embeddings is introduced that captures several complementary properties required for downstream effectiveness, demonstrating that table-level embedding quality cannot be reduced to retrieval alone.
Abstract
Tabular data is the dominant structured-data modality, and learning table representations has become a core research direction. Table-level embeddings in particular underpin a wide range of applications, including table retrieval, data lake discovery, and table classification. Despite their importance, there is still limited understanding of how different embedding approaches behave across tasks, making systematic evaluation and analysis essential. In this work, we introduce a systematic evaluation of table-level embeddings that captures several complementary properties required for downstream effectiveness. We realize this evaluation by extending TEmBed, a recently proposed testbed for tabular embeddings, whose table-level coverage is currently limited to a single retrieval task. An empirical study over the TEmBed model pool confirms that no single model excels across all tasks, demonstrating that table-level embedding quality cannot be reduced to retrieval alone.
An initial approach that first aligns heterogeneous table-cell representations into a shared space using Hirschfeld–Gebelein–Rényi maximal correlation (HGR) is proposed, and it is found that it generalizes competitively compared to models specifically designed for individual tasks.
Tabular in-context learners such as TabPFN, Mitra, or ConTextTab rely on alternating row and column attention over 2D sequences of latent embeddings. These attention patterns differ markedly from the one-dimensional case in language models: row attention involves longer sequences while column attention operates on much...
Maximilian Schambach, Clemens Biehl, Sam Thelin· 0 citations
Dataset search often aims to identify joinable or unionable datasets to augment a given query table. State-of-the-art approaches rely on large language models (LLMs) to embed tables into vector representations and perform semantic similarity search. However, existing work assumes a centralized data repository with em...
Lennart Behme, Emil Badura, Leonard Geißler et al.· Proceedings of the VLDB Endo...· 0 citations
This work introduces DocHop, a benchmark for integrated chart--context reasoning in document-style images and constructs DocHop via a stochastic logic-first generation pipeline with controllable reasoning depth and visual density, to enable systematic evaluation.
Zhuoran Yu, Le Thien Phuc Nguyen, Jaden Park et al.· 1 citation
LSR retains a general encoder as a transferable semantic substrate while explicit modules reorganize, augment, or index its representations for a defined operational objective.
Pedro Emílio Amador Salomão· Nexus Science Review· 0 citations
This work proposes H2Table (Hierarchical Hypergraph-Enhanced Table Reasoning), a novel framework that represents complex tables as hierarchical nested hypergraphs, and designs a tailored hypergraph encoder to facilitate message passing between hyperedges and nodes within complex tables.
Jia Ling, Yang-Fan Wang, Chen Tang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.