Skip to content

Solving In-Table Prediction Problems by Deep Neural Networks with Performance Evaluation Using Synthetic Data

Sep 2026 · 0 citations · 38 references
Computer Science

TL;DR

This work investigates whether NNs can predict the values of arbitrarily selected columns in a given table based on the remaining known columns, and concludes that the attention-based structure outperforms the other two networks, when a sufficiently large number of training examples is available and a relatively large embedding length is chosen.

Abstract

Tabular deep learning (TDL) leverages neural networks (NN) to extract patterns from tabular data. Traditional TDL methods follow a supervised learning paradigm, where a target feature is explicitly given. In this work, however, we explore a different approach by employing deep NNs to learn relationships among individual columns within a given table. We investigate whether NNs can predict the values of arbitrarily selected columns in a given table based on the remaining known columns. We call this problem In-Table Prediction (ITB), which is slightly different from table imputation methods and the pretraining task of TDL. Three potential usage scenarios are identified, which, to our best knowledge, have not been extensively studied in the literature. A self-supervised learning approach is applied to address this problem by randomly selecting columns to be masked out and used as learning targets. This work focuses on tabular datasets containing only continuous features. To handle missing values in continuous features, a novel neural layer is proposed to embed both numerical and empty values. Synthetic data is generated based on predefined column relationships, with empty values inserted using two distinct mechanisms. Additionally, an adapted masking strategy is employed to create test data. Performances of three NN architectures, namely MLP, Resnet and Transformer, are evaluated using the generated synthetic data. We conclude that, the attention-based structure outperforms the other two networks, when a sufficiently large number of training examples is available and a relatively large embedding length is chosen. We stress that these findings are obtained under controlled, synthetic conditions with a small number of columns and it should therefore be regarded as an initial, narrowly-scoped investigation rather than a general characterization of ITP on real-world tabular data.

View source

Similar papers

Open access Aug 2026

In-Context Learning Meets Small Molecule Property Prediction: Benchmarking Novel Machine Learning Approaches.

Recently, a new category of machine learning approaches for tabular data has emerged: tabular foundation models (TFM), based on in-context learning. A TFM is a neural network (usually a transformer) pretrained primarily on synthetic data. Its input is an entire data set: features and labels for training records, along...

D. Matyushin, A. Sholokhova · 0 citations
Conference Aug 2026

XAI-AutoDL: An Explainable Automated System for Deep Learning-Based Classification Using Meta-Features

The increasing use of machine learning (AutoML) and deep learning means hyperparameter optimization can save time. Nevertheless, it is not clear how they compare on structured tabular classification under tight computational constraints. Here we present a resource-aware benchmark that compares the performance of Random...

Priya Surana, Palash Hemade, Ishan Patil et al. · 0 citations
#machine learning Preprint Oct 2026

Compact set-valued deep ensembling in multi-class classification

This paper tackles visible challenges in deep ensemble learning, where deep neural networks serve as ensemble members: training and storage burdens, and robustness of cautious (set-valued) predictions targeting multiple utilities, which may involve reward-sensitivity. To mitigate the training and storage burdens, we pr...

Kim-Dung Tran, D. Nguyen, Vu-Linh Nguyen et al. · 0 citations
Review

Learning Strategies with Attentive Neural Processes

This work proposes a novel LAL method for classification that exploits symmetry and independence properties of the active learning problem with an Attentive Conditional Neural Process model and gives the model the ability to adapt to non-standard objectives.

Tim Bakker, H. van Hoof, Max Welling · 0 citations
Preprint Aug 2026

When Does Self-Supervised Pretraining Help Tabular Models? A Study of Label Scarcity and Missing Data

While SSL outperforms training from scratch on average and remains competitive with state-of-the-art tree ensembles, the SSL-vs-scratch gains exhibit high inter-task variance and lack significance, indicating the findings reflect general properties of tabular SSL rather than idiosyncrasies of one particular pretext tas...

Sahand Mazrouei · 0 citations
Open access Aug 2026

Benchmarking Pre-Trained Feature Extractors: A Comparative Study Across Deep Learning Tasks

This study systematically compares seven pre-trained feature extractors across three architectural families, convolutional neural networks (CNNs), Vision Transformers (ViTs), and self-supervised models to provide practical guidance on model selection for downstream deep learning tasks.

Rafeek Sibrikhan, M. Mufassirin · 0 citations

Related blog posts

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.