Skip to content

Kolmogorov-Arnold Networks for Small Language Models

Jul 2026 · arXiv.org · Vol abs/2607.15525 · 0 citations · 31 references
Computer Science

TL;DR

Small-basis KANs provide a practical, corpus-transferable interface for auditing learned scalar transformations, but the tested replacements show no consistent benchmark, quality, or latency advantage over strong MLP baselines.

Abstract

Kolmogorov--Arnold Networks (KANs) replace fixed node activations with learned one-dimensional edge functions, offering an explicit interface for interpretation and a possible alternative to transformer feed-forward networks. We test these claims separately. In a six-layer, 10M-parameter B-spline KAN, we reconstruct all 884,736 feed-forward edges: 87.8\% exceed (NLS>0.1) and 0.4\% are inactive. Pruning the lowest-activity 20--25\% causes negligible loss increase, although structured MLP neuron pruning tolerates comparable sparsity. The audit replicates on BabyLM, but grid-size sweeps show that near-total fPCA compression and high closed-form-fit coverage are properties of the low-capacity grid-2 basis, not universal KAN behavior. For replacement, we evaluate MLP, SwiGLU, grouped Chebyshev, and rational GR-KAN networks on BabyLM. The KAN-family and gated variants improve validation loss over the GELU MLP, but this ordering does not transfer to standardized benchmarks: across ten seeds and 59,875 BLiMP pairs, accuracies span 62.4--63.1\%, EWoK remains at chance, and a (+0.7)-point GR-KAN effect on BLiMP reverses on the supplement. Larger tests are also cautionary: parameter-matched MLPEdge underperforms the MLP on Wikitext-103, and 286M-parameter GR-KAN remains below a SwiGLU ClimbMix baseline after stabilization. Thus, small-basis KANs provide a practical, corpus-transferable interface for auditing learned scalar transformations, but the tested replacements show no consistent benchmark, quality, or latency advantage over strong MLP baselines.

View source

Similar papers

Preprint Aug 2026

SparseKAN: Compressing Kolmogorov--Arnold Networks Across Basis Functions, Neurons, and Bits

SparseKAN is presented, a unified approach that compresses KANs along three complementary axes: basis functions, neurons/channels, and numerical precision, and demonstrates that SparseKAN converts functional redundancy into measurable software and hardware efficiency.

Kazi Ahmed Asif Fuad, Lizhong Chen · 0 citations
Preprint Aug 2026

ER-KANs: Efficient and Robust Kolmogorov-Arnold Networks for Data-Scarce Scientific Machine Learning

The efficient-KAN literature---covering Chebyshev, wavelet, and radial-basis-function variants of the original Kolmogorov-Arnold Network---has been benchmarked almost entirely on clean data. We show that this choice conceals a large capability difference between architectures: ChebyKAN's test MSE (evaluated against cle...

Harshil Lodhiya · 0 citations
Preprint Aug 2026

An Embedded RISC-V Evaluation of Kolmogorov--Arnold Networks in Hard-Constrained Recurrent Physics-Informed Models

Hard-constrained recurrent physics-informed networks (HRPINNs) embed known dynamics inside a recurrent numerical integrator and restrict a neural branch to learning only the residual dynamics that the first-principles model does not capture. Kolmogorov--Arnold Networks (KANs) have been proposed as parameter-efficient r...

Enzo Nicolás Spotorno, J. Leal · 0 citations
Preprint Aug 2026

HYDRA: Hyperbolic Dynamic Representation Architecture for Kolmogorov-Arnold Networks

HYDRA is a parameter-efficient hyperbolic extension of KAN that combines spline-based functional learning with representations in the Poincar\'e ball, and consistently achieves competitive or superior predictive performance while improving parameter efficiency and representation interpretability.

Zhao Su, Yuxin Xia, Haoran Li et al. · 0 citations
Open access Aug 2026

KAN-Payne: A Controlled Evaluation of Kolmogorov–Arnold Networks for Stellar Spectral Emulation and Label Recovery

Kolmogorov–Arnold networks (KANs) replace the fixed activations of a multilayer perceptron (MLP) with learnable univariate edge functions. We evaluate KANs as Payne-style label-to-flux emulators using the public 1000-spectrum Kurucz grid released with The Payne. Two capacity-matched KAN–MLP pairs (approximately 2.3 and...

Shuo Zhang, Rui Wang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.