Skip to content

All you need is SAMPAT

Jul 2026 · arXiv.org · Vol abs/2607.09235 · 0 citations · 8 references
Computer Science Mathematics

TL;DR

A three layer neural architecture, SAMPAT (Smooth Approximation via Multivariate Polynomials and Analytic Transformations), that can provably learn a continuous, everywhere differentiable function, that can approximate any smooth function arbitrarily closely is presented.

Abstract

The current state of the art in AI/ML rests on deep neural architectures, which, in general, suffer from a lack of interpretability. Interpretability is crucial to gleaning insights while analyzing experimental data, where quantitative predictions may not be adequate for a scientist. We present a three layer neural architecture, SAMPAT (Smooth Approximation via Multivariate Polynomials and Analytic Transformations), that can provably learn a continuous, everywhere differentiable function, that can approximate any smooth function arbitrarily closely. SAMPAT's approximant can be expressed as a closed and compact algebraic, analytic expression, providing complete interpretability. Experiments on synthetic and benchmark datasets indicate that SAMPAT yields competitive performance with simpler representations. For many tasks, a two layer SAMPAT suffices. By imposing restrictions on the connectivity between neurons, SAMPAT may be used to provide a range of approximants, including regular and trigonometric polynomials, rational expressions, Gaussians, mixtures of Gaussians, as well as arbitrary combinations of the same; without restrictions, it learns a suitable structure. SAMPAT may be used to factorize polynomials and model nonlinear systems. With the addition of skip connections, a 4 to 6 layer SAMPAT is adequate to represent a substantive range of methods widely used in AI/ML, allowing the choice of a model's family, not just its parameters, to also be optimized as part of the learning process.

View source

Similar papers

Review Aug 2026

Not All Attention Is Equal: A Quantitative Survey of the EEI Trade-off

This survey traces attention from Bahdanau-Luong alignment through the Transformer and into vision architectures, and reviews fixed and learned sparse attention, linear attention, IO-aware exact algorithms including FlashAttention, and state-space alternatives including Mamba.

Aditya Singh · 0 citations
Preprint Aug 2026

The Sparsity Whisperer

A family of difference-informed pruning methods built upon this principle are introduced, suggesting that preserving output differences is a broadly useful and composable signal for post-training LLM sparsification.

Linghao Kong, Inimai Subramanian, Micah Adler et al. · 0 citations
Preprint Aug 2026

On the Principles Behind Neural Network Optimizers

This thesis develops a principled grounding for Adam and motivates new designs, and reveals new local structures in matrix-based nonconvex problems, and helps understand and improve recent NN optimizers, such as Muon.

Yu-Shun Zhang · 0 citations
#artificial intelligence Preprint Sep 2026

Exactness at Inference: A Representational Criterion for Out-of-Distribution Generalization

A model generalizes outside its training distribution only when it computes a representation structurally equivalent to the generating mechanism, not an approximation fitted to it. Such equivalence is necessary for exactness in and out of distribution, and extrapolation is governed by this exactness at inference, whate...

Filipe Marinho Rocha, I. Dutra, V. Costa et al. · 0 citations
Preprint Aug 2026

TESLA: Taylor Expansion of Sinusoidal Learnable Activations

TESLA, an activation defined as a learnable combination of sine and cosine terms, enabling explicit control over polynomial degree and selective amplification of high-order components is proposed, indicating that activation-level degree control transfers to more general vision workloads.

Daehwa Ko, Jae-Hwan Kim, Seunghyun Ham et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Relative Generalization Invariance of LLM Pretraining

This work introduces Relative Generalization Invariance (RGI), the invariance of the validation-loss difference between any two tokens across models, and shows that RGI approximately holds across a wide range of optimizers and moderate architectural variations, suggesting that these choices induce an approximately unif...

Feng-Zhuo Zhang, Shu-Che Wang, Sheng-Gui Li et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.