Aug 2026· SIAM Journal on Scientific Computing· Vol 48, pp. 1011-· 0 citations· 9 references
Computer Science
TL;DR
Results suggest, and experimental evidence corroborates, that kernel machines relying on empirical kernels extracted from trained DNNs can act as surrogates for trained finite-width DNNs, and determine regimes where both kernel machines trained on features extracted from an underlying DNN are demonstrably superior to the latter.
Abstract
Abstract.
The impressive performance of deep neural networks (DNNs) on a variety of learning tasks has spurred much investigation into improving their training and characterizing the trained functions. Recent work has shown the equivalence of DNNs in the infinite width limit and kernel machines relying on the neural tangent kernel (NTK) at initialization. These results suggest, and experimental evidence corroborates, that kernel machines relying on empirical kernels extracted from trained DNNs can act as surrogates for trained finite-width DNNs. The high computational cost of assembling the NTK, however, makes this approach infeasible in practice. In the current work, we study the performance of the conjugate kernel (CK), an efficient approximation to the NTK. For smooth function and logistic regression, we show that the CK performance is only marginally worse than that of the NTK and, in certain cases, much more superior. In particular, we establish bounds for the test losses, verify them with numerical tests, and identify the regularity of the kernel as the key determinant of performance. We also determine regimes where both kernel machines trained on features extracted from an underlying DNN are demonstrably superior to the latter and use this to suggest a recipe for accelerating DNN performance inexpensively. We present a demonstration of this on foundation models by comparing their performance on a classification task using a conventional technique and our prescription. We also show how our approach can be used to improve physics-informed operator network training as well as convolutional neural network training for vision classification tasks.
This work proposes an NC-inspired training framework for simplifying deep networks during training, monitoring representation dynamics through the Inverse Fisher Criterion to identify both the split point between feature extraction and classification and the training stage at which simplification becomes viable.
Lorenzo Sciandra, Samuele Fonio, Roberto Esposito· arXiv.org· 0 citations
Kernel methods are widely used because of their strong theoretical guarantees and empirical performance. However, their high computational cost limits their applicability to large-scale datasets. To address this shortcoming, several approaches use Maximum Mean Discrepancy to construct representative subsets that preser...
Blanca Cano-Camarero, Yago R. Aguado-Carrillo-de-Albornoz, Ángela Fernández-Pascual et al.· 0 citations
This work introduces a general kernel-based encoder-decoder framework for operator learning that separates observation, representation, learning, and reconstruction, and develops this framework for multi-input, multi-output operator learning, where operators map between products of potentially distinct function spaces.
Adrien Weihs, Chun-Yang Liao, Jingmin Sun et al.· 0 citations
This paper deeply integrates convex optimization theory with the backpropagation algorithm and constructs a novel stable and efficient training mechanism for neural networks that achieves favorable adaptability to both shallow fully connected networks and deep convolutional networks.
Weiwei Guo· Applied and Computational En...· 0 citations
Stochastic weight averaging applied to classification as an alternative ensembling technique that does not require repeated training runs and provides an equivariance boost that goes beyond what could be expected from the performance increase due to SWA alone.
Longde Huang, Axel Flinth, Jan E. Gerken· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.