Skip to content
Preprint

On the Principles Behind Neural Network Optimizers

Aug 2026 · 0 citations · 60 references
Computer Science Mathematics

TL;DR

This thesis develops a principled grounding for Adam and motivates new designs, and reveals new local structures in matrix-based nonconvex problems, and helps understand and improve recent NN optimizers, such as Muon.

Abstract

Reliable optimization is central to neural network (NN) training, yet Adam, the default optimizer for modern LLMs, rests on a fragile foundation. This thesis develops a principled grounding for Adam and motivates new designs. First, we revisit Adam's divergence--convergence debate and show the existence of a problem-dependent phase transition: with properly chosen, batch-size-dependent hyperparameters, Adam converges, whereas under small-$\beta_2$ regimes it can diverge. Second, we investigate why Adam substantially outperforms SGD on Transformers through Hessian structure. We find that the Hessian evolves toward a near-block-diagonal form along training, accompanied by strong block heterogeneity. We prove that this structure makes Adam's diagonal preconditioner effective. We further show that this special Hessian structure originates from consecutive multiplications of large matrix variables, and we provide a rigorous analysis based on random matrix theory. Finally, these insights motivate Adam-mini, a new optimizer that reduces Adam's memory footprint by 50\% while preserving its performance. Our results also have broader implications beyond Adam: they reveal new local structures in matrix-based nonconvex problems, and also help understand and improve recent NN optimizers, such as Muon.

View source

Similar papers

Preprint Aug 2026

The Loss Does Not See the Basis, but Adam Does

Gradient descent on a factored model $W = UV^\top$ is implicitly biased toward low-rank solutions, while Adam, starting from the same small initialization, is not. We trace the difference to the gauge symmetry of the loss, its invariance under $(U, V) \mapsto (UQ, VQ)$. Gradient flow's low-rank mechanism is available t...

D. Singh · 0 citations
Jul 2026

Between Gradient and Natural Gradient: A Continuum of LoRA Initializations

It is shown that a principled design space for LoRA initialization and curvature preconditioning should be treated as a tunable dimension rather than a fixed design decision, and a deployable, search-free variant, ULoRA-Auto, selects per-layer exponents from measured spectral statistics, approaches this upper bound at...

Dian Liu, Farshid Ghezelbash · 0 citations
#machine learning Preprint Aug 2026

Convergence rates for the RMSprop optimizer with full control of the hyperparameters

The key innovative new feature in the proof of the analysis are suitable inverse moment bounds for the second moment process in RMSprop that hold not just for all sufficiently large n but hold for every gradient step $n=1,2,3,...$ with all error constants being explicitly specified.

Steffen Dereich, Arnulf Jentzen · 0 citations
Open access Sep 2026

NeuroFuzzyAdam: A Fuzzy-Enhanced Adaptive Optimization Algorithm for Deep Learning

Deep neural networks are commonly trained using adaptive optimization methods such as Adam because they converge quickly and perform well under stochastic training conditions. Despite these advantages, recent research has revealed several important drawbacks of Adam. In particular, the optimizer tends to converge towar...

Susilo Hariyanto, Siti Khabibah, Retno Putri Dwi Rahmawati et al. · 0 citations
Open access Aug 2026

Dynamic Regime Maps of Neural Network Training Under the Adam Optimizer: An Observable-Based Empirical Analysis

Adaptive optimization algorithms are fundamental to modern deep learning; however, the global organization of neural-network training regimes induced by optimizer hyperparameters remains insufficiently understood. In particular, the influence of the Adam moment coefficients on the stability and qualitative behavior...

S. Sveleba, I. Katerynchuk, I. Kunyo et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.