Skip to content
Preprint

Geometry-Constrained Kolmogorov-Arnold Networks: Learning Edge Geometry via Banach Duality

Aug 2026 · 0 citations · 25 references
Computer Science Mathematics

TL;DR

Geometry-constrained KANs are introduced, a family of edge activations derived from Banach duality maps in which the geometry itself is learned through a scalar exponent $p>1$ per edge, which controls the qualitative response.

Abstract

Kolmogorov-Arnold Networks (KANs) replace fixed activations in deep architectures with learnable univariate edge functions, making the choice of edge parametrisation central. Existing variants rely on fixed bases such as splines, polynomials, or Fourier features, which impose a function-space geometry before data are observed. We introduce geometry-constrained KANs, a family of edge activations derived from Banach duality maps in which the geometry itself is learned through a scalar exponent $p>1$ per edge. This exponent controls the qualitative response: sub-Euclidean values produce sharp, threshold-like behaviour reminiscent of the $\ell_1$ (LASSO) geometry, $p = 2$ recovers the linear regime, and larger values produce flatter responses near the origin. Across 50 symbolic-regression targets ($40$ from the AI Feynman benchmark plus $10$ synthetic stress tests), geometry-constrained KANs match or beat every fixed-basis baseline on median NRMSE (Banach-KAN $0.030$, tying Chebyshev and improving on splines); on average rank Banach-KAN is best on the $18$-equation core ($2.00$) and statistically tied with the strongest spline on the full benchmark ($2.32$ vs. $2.34$). The clearest gains appear under measurement noise: as $\sigma$ grows from $0$ to $1$, $\ell^p$-KAN degrades only $3.7\times$ -- below even a cross-validated spline ($\approx 11\times$) -- while an unregularised spline degrades $21.6\times$; Banach-KAN degrades $8.8\times$, comparable to a tuned spline but far more stable than the unregularised one. Banach-KAN also takes the most per-equation wins in the small-sample regime, with fixed-basis models catching up only as the training set grows. Learned exponents provide an interpretable, relative signal: at a fixed initialisation they reveal a consistent, target-dependent geometric ordering across equation families and input dimensions.

View source

Similar papers

Preprint Aug 2026

HYDRA: Hyperbolic Dynamic Representation Architecture for Kolmogorov-Arnold Networks

Kolmogorov-Arnold Networks (KANs) enhance nonlinear function approximation by replacing scalar weights with learnable univariate functions. However, assigning an independent function to every connection results in substantial parameter redundancy, limiting their scalability and efficiency. To reduce this redundancy, we...

Zhao Su, Yuxin Xia, Haoran Li et al. · 0 citations
Preprint Aug 2026

SparseKAN: Compressing Kolmogorov--Arnold Networks Across Basis Functions, Neurons, and Bits

SparseKAN is presented, a unified approach that compresses KANs along three complementary axes: basis functions, neurons/channels, and numerical precision, and demonstrates that SparseKAN converts functional redundancy into measurable software and hardware efficiency.

Kazi Ahmed Asif Fuad, Lizhong Chen · 0 citations
Preprint Aug 2026

Sharp Sobolev Approximation on General Domains by Linearized Shallow Networks with Analytic Activations

It is proved that quasi-Chebyshev parameter sets with univariate resolution $m$ generate fixed feature spaces attaining the sharp $H^r$-to-to-H^s$ approximation order for a class of analytic activations satisfying a quantitative non-cancellation condition on their Taylor coefficients.

Jia Li, Tong Mao, Jin-Chao Xu · 1 citation
#artificial intelligence Preprint Sep 2026

RecKAN: Kolmogorov-Arnold Networks with a Learnable Recursive Polynomial Basis

Kolmogorov--Arnold Networks (KANs) replace the fixed scalar weights of a standard network with learnable univariate functions on each edge, but existing variants still fix the \emph{basis} that those functions are built from: B-splines, Chebyshev polynomials, wavelets, or Jacobi polynomials, and learn only the combinat...

Amirhosein Azarpour · 0 citations
Preprint Aug 2026

Operator-Theoretic Generalization Bounds for Multitask Deep Learning

Operator-theoretic generalization bounds for deep multi-output function classes are developed by representing network layers as Koopman composition operators on vector-valued reproducing kernel Hilbert spaces and derive Rademacher complexity bounds for invertible and width-expanding injective architectures.

Mahdi Mohammadigohari, Thomas Borsani, G. Di Fatta · 1 citation
Preprint Aug 2026

Fixed and Adaptive Topological DeepONets: Functional Measurements on Hausdorff Locally Convex Spaces

The Topological DeepONet framework is built on, replacing point samples by continuous linear functionals drawn from the continuous dual of a Hausdorff locally convex space, whose topology is generated by a point-separating family of seminorms rather than a single norm, and develops fixed and adaptive functional measure...

Khemraj Shukla, G. Karniadakis · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.