Stochastic weight averaging applied to classification as an alternative ensembling technique that does not require repeated training runs and provides an equivariance boost that goes beyond what could be expected from the performance increase due to SWA alone.
Abstract
The symmetries of a learning task have become an important factor in designing modern deep learning solutions. Data augmentation is a straightforward and effective way of incorporating symmetries into a generic neural network. Recent results show that infinitely large deep ensembles show perfect symmetry when trained on augmented data. However, since training ensembles requires repeating the training process many times, this method is costly. In this work, we study stochastic weight averaging (SWA) applied to classification as an alternative ensembling technique that does not require repeated training runs. We analyze SWA by approximating the stochastic training trajectory at the end of training with an Ornstein--Uhlenbeck process. We show that in the infinite-width limit, SWA on augmented data provides an equivariance boost that goes beyond what could be expected from the performance increase due to SWA alone. We verify our results with extensive numerical experiments on numerous models spanning image and graph classification with both discrete and continuous symmetries.
This work proposes an NC-inspired training framework for simplifying deep networks during training, monitoring representation dynamics through the Inverse Fisher Criterion to identify both the split point between feature extraction and classification and the training stage at which simplification becomes viable.
Lorenzo Sciandra, Samuele Fonio, Roberto Esposito· arXiv.org· 0 citations
An exact finite-sample necessary-and-sufficient condition for two-stage training to strictly outperform real-only training under the same real-data budget and identical real-stage updates is established.
A single taxonomy of the new regularization methods such as adaptive regularization, information-theoretic constraints, structured sparsity, stochastic regularization and regularization at the representation level is presented and Hybrid Adaptive Information Regularization (HAIR) is suggested which is a dynamic complex...
Rak esh, A. An· International Journal of Mac...· 0 citations
Results suggest, and experimental evidence corroborates, that kernel machines relying on empirical kernels extracted from trained DNNs can act as surrogates for trained finite-width DNNs, and determine regimes where both kernel machines trained on features extracted from an underlying DNN are demonstrably superior to t...
S. Qadeer, A. Engel, Amanda A. Howard et al.· SIAM Journal on Scientific C...· 0 citations
This work presents a method to accelerate the optimization of learning high dimensional functions using deep neural network (DNN), and studies the effect of adding features which distill pretrained DNN into TNs using a discretize and decompose strategy.
The Broad Learning System (BLS) enables efficient closed-form training by solving the output weights in a single ridge-regression step. However, this one-shot full-batch estimator often exhibits a pronounced generalization gap. This paper proposes Batch Weight-Averaged BLS (BWA-BLS), which partitions the training set i...
Chen Xu· International Conference on...· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduOct 2, 2026
Martin Trust Center Managing Director Bill Aulet introduces Dear Dreamer, a free platform for middle and high school students who want to learn about entrepreneurship.
Microsoft Research Blog· microsoft.comSep 30, 2026
Extreme space-weather events can damage power systems on Earth and degrade GPS accuracy and satellite operations. A new machine learning system can predict where damage is likely to occur 30-60 minutes before a storm arrives. The post Forecasting space weather risks on power grids appeared first on Microsoft Research.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.