Skip to content
Preprint

Tabular Numeric Stretch Transformation

Aug 2026 · 0 citations · 36 references
Computer Science

TL;DR

The stretch transformation framework is introduced, which formulates numeric feature preprocessing as an optimization problem to make the target function smoother and thus more learnable and shows that explicitly optimizing for target function smoothness is a powerful and underexplored strategy for tabular deep learning.

Abstract

Tabular data presents unique challenges for deep learning due to its heterogeneous nature, where numeric features exhibit diverse distributions, scales, and statistical properties. Although recent advances have improved how models learn from tabular data, how numeric data are transformed into model-friendly representations remains comparatively underexplored. We introduce the stretch transformation framework, which formulates numeric feature preprocessing as an optimization problem to make the target function smoother and thus more learnable. Our framework has two variants: (1) unsupervised stretch, which uniformly redistributes feature density via minimax optimization, and (2) supervised stretch, which optimizes target-aware numeric feature transformations from the perspective of target-function smoothness by minimizing the target function's Dirichlet energy in the transformed space. Our theoretical analysis further connects this framework to several popular transformations: unsupervised stretch is closely related to Piecewise Linear Encoding through a shared piecewise-linear geometry and approaches the empirical CDF transformation as the number of bins grows, while supervised stretch becomes closely related to target encoding in the fine-binning limit. Comprehensive experiments on 38 datasets from the TALENT benchmark demonstrate that supervised stretch consistently outperforms all baselines. These results show that explicitly optimizing for target function smoothness is a powerful and underexplored strategy for tabular deep learning.

View source

Similar papers

#machine learning Preprint Aug 2026

TabNSM: Neural Sparse Mixer for Tabular Regression

TabNSM provides an effective and scalable approach to deep tabular regression, and demonstrates that selective interaction modeling, structured regression supervision, and difficulty-aware sampling provide an effective and scalable approach to deep tabular regression.

Ali Eslamian, Qiang Cheng · 0 citations

FeatureZ : A General Framework for Feature-Preserving Compression via Pointwise Bounds and Star Classification

This paper introduces FeatureZ, a lossy compression framework for structured volumetric scalar fields that can preserve a wide class of geometric and topological features and demonstrates that FeatureZ preserves diverse features during compression with minimal overhead.

Nathaniel Gorski, Xin Liang, Han-Qi Guo et al. · 0 citations
Book Open access Aug 2026

Efficient Piecewise-Linear Embeddings for Deep Tabular Regression by Guided Breakpoint Allocation

GBDT-Guided Piecewise-Linear (GGPL) embeddings are introduced, which leverage GBDTs to provide a strong data-driven prior for breakpoint placement and fine-tune these breakpoints via gradient descent, which establishes GGPL as an effective numerical embedding for deep tabular regression.

Min-Kook Suh, M. Eo, Kyungeun Lee et al. · 0 citations
#machine learning Preprint Sep 2026

Neural Symbollic Regression Using Deep Learning and Sparse Modelling

A scalable and understandable neural-symbolic framework that treats neural networks as functional preconditioners for symbolic discovery, creating a solid link between neural approximation and the discovery of sparse equations for scientific machine learning.

U. Ravikumar, S. Sumitra · 0 citations
Preprint Aug 2026

Across the Loss Landscape with Progressive Growth

Under standard local regularity conditions around non-degenerate minima, it is proved that local sublevel sets are well approximated by ellipsoids and that basin accessibility under frozen constraints can be characterized by an explicit effective curvature in the frozen directions.

Paul Caillon, Christophe Cerisara, Alexandre Allauzen · 0 citations
#machine learning Preprint Sep 2026

SMILE: Bridging Continuous Optimization and Discrete Symbolic Recovery

Symbolic regression (SR) discovers closed-form mathematical expressions from data, offering interpretability beyond black-box models. Existing methods suffer from slow convergence in combinatorial search spaces and lack mechanisms to exploit compositional structure in the data. We introduce SMILE (Sine, Multiplication,...

Mansooreh Montazerin, Antonio Ortega, Ajitesh Srivastava · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.