Skip to content
Preprint

When Does Self-Supervised Pretraining Help Tabular Models? A Study of Label Scarcity and Missing Data

Aug 2026 · 0 citations · 27 references
Computer Science

TL;DR

While SSL outperforms training from scratch on average and remains competitive with state-of-the-art tree ensembles, the SSL-vs-scratch gains exhibit high inter-task variance and lack significance, indicating the findings reflect general properties of tabular SSL rather than idiosyncrasies of one particular pretext task.

Abstract

Self-supervised learning (SSL) has emerged as a promising approach for tabular data, yet its efficacy under extreme label scarcity and test-time missingness remains under-explored. In this paper, we evaluate a mask-and-recover SSL pretraining objective against training from scratch and classical baselines across 14 diverse classification tasks. First, while SSL outperforms training from scratch on average and remains competitive with state-of-the-art tree ensembles (achieving ~0.8954 AUC vs. Random Forest's 0.9015 at 10% labels), the SSL-vs-scratch gains exhibit high inter-task variance and lack significance (p = 0.626 at both 5% and 10% labels). Second, contrary to the hypothesis that missing-value imputation objectives universally benefit datasets with native missingness, SSL yields the most reliable improvements on clean datasets, while frequently degrading performance on datasets with high inherent missingness. Third, despite this training variance, SSL-pretrained models achieve a higher average AUC than scratch-trained models under both test-time missingness completely at random (MCAR) injection (+0.0245 AUC, positive on 11 of 14 tasks) and structured missingness shifts (MNAR, +0.0418 AUC, positive on 8 of 14 tasks), though neither difference remains statistically significant after Holm-Bonferroni correction for multiple comparisons (adjusted p = 0.118 and p = 0.518, respectively). Fourth, comparing our mask-and-recover objective against three established tabular SSL baselines (VIME, SCARF, SubTab) under an identical encoder architecture, we find no significant difference from any of them (adjusted p = 0.459, p = 1.000, p = 1.000), indicating our findings reflect general properties of tabular SSL rather than idiosyncrasies of one particular pretext task.

View source

Similar papers

Preprint Aug 2026

C-Score: Beyond Accuracy for Robustness Assessment in Semi-Supervised Learning under Open-World Unlabeled Contamination

C-Score, a compact framework that evaluates training behavior in three complementary spaces: prediction, feature representation, and optimization, suggests that clean accuracy alone is insufficient for evaluating SSL robustness in open-world environments, and that internal diagnostic signals are necessary for more reli...

Tsao-Lun Chen, Chicheng Fu, Han-Yi Chou et al. · 0 citations
Jul 2026

Boosting Semi-Supervised Learning With Entropy-Guided Adaptive Reward Maximization

Existing semi-supervised learning (SSL) methods rely predominantly on pseudo-labeling and consistency regularization to leverage unlabeled data, demonstrating significant performance improvements. However, we pinpoint that these methods suffer from a confidence-for-weighting issue, overvaluing high-confidence pseudo-la...

Anyang Tong, Zenglin Shi, Zhun Zhong et al. · 0 citations
Jul 2026

Robust Ensemble Learning Under Label Noise: A Theoretical Analysis and Framework-Specific Solutions.

This article utilizes bias-variance-diversity (BVD) decomposition theory to examine the impact of noisy labels on three mainstream ensemble paradigms: Bagging, Boosting, and Stacking and characterizes the mechanisms behind performance degradation.

Guanxiong He, Jie Wang, Zhiyong Li et al. · 1 citation
#artificial intelligence Preprint Aug 2026

Forget or Fine-tune? A Comparative Study of Machine Unlearning Strategies for Noisy Label Correction

This work conducts a comparative empirical study of five MU methods across symmetric, asymmetric, instance-dependent, and open-set noise on CIFAR-10, CIFAR-100, and the real-world noisy dataset Food-101N and finds that the appropriate unlearning strategy is conditioned on the noise structure.

J. L. Sant'Ana, Filipe R. Cordeiro · 0 citations
Review Open access Aug 2026

Conformal prediction for multi-label learning: a review of methods and guarantees.

This review consolidates the landscape of CP adaptations for MLL under a unified framework, examining the types of outputs and guarantees they provide, where label dependencies are incorporated, and how inference cost scales with the number of labels.

Harris Papadopoulos · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.