Skip to content

Effects of width-dependent model hyperparameters and ℓ2-regularization on the loss landscape of two-layer ReLU networks

Jul 2026 · arXiv.org · Vol abs/2607.16720 · 0 citations · 42 references
Computer Science

TL;DR

Interestingly, the experiments reveal that using AdamW as an optimizer prevents the collapse of the learned parameters, whereas using SGD does not, which may help explain the success of AdamW in deep learning training.

Abstract

Understanding deep neural networks remains a central challenge in machine learning. In particular, the theoretical properties of even two-layer ReLU networks, especially in the presence of weight decay, remain poorly understood. To this end, we derive a sufficient condition on the hyperparameter settings under which the global minima collapse to the zero solution. Interestingly, our experiments reveal that using AdamW as an optimizer prevents the collapse of the learned parameters, whereas using SGD does not, which may help explain the success of AdamW in deep learning training. In addition, when restricting the input dimension to one, we derive an analytical solution for the globally optimal parameter sets of two-layer ReLU networks and show that $\ell_2$-regularization has a width-invariant effect on connectivity, but its dimensionality-reducing effect becomes stronger as the network width increases. These results provide insight into how width-dependent hyperparameters influence the geometry of regularized loss landscapes.

View source

Similar papers

Open access 2026

Novel Regularization Methods to Prevent Overfitting in Machine Learning Models

A single taxonomy of the new regularization methods such as adaptive regularization, information-theoretic constraints, structured sparsity, stochastic regularization and regularization at the representation level is presented and Hybrid Adaptive Information Regularization (HAIR) is suggested which is a dynamic complex...

Rak esh, A. An · 0 citations
Open access Aug 2026

Dynamic Regime Maps of Neural Network Training Under the Adam Optimizer: An Observable-Based Empirical Analysis

Adaptive optimization algorithms are fundamental to modern deep learning; however, the global organization of neural-network training regimes induced by optimizer hyperparameters remains insufficiently understood. In particular, the influence of the Adam moment coefficients on the stability and qualitative behavior...

S. Sveleba, I. Katerynchuk, I. Kunyo et al. · 0 citations
Preprint Aug 2026

On the Principles Behind Neural Network Optimizers

This thesis develops a principled grounding for Adam and motivates new designs, and reveals new local structures in matrix-based nonconvex problems, and helps understand and improve recent NN optimizers, such as Muon.

Yu-Shun Zhang · 0 citations
Preprint Aug 2026

Accelerated Learning of High Dimensional Functions with a Tensor-Featured Training Network

This work presents a method to accelerate the optimization of learning high dimensional functions using deep neural network (DNN), and studies the effect of adding features which distill pretrained DNN into TNs using a discretize and decompose strategy.

Karl Pierce, Y. Khoo, Haizhao Yang · 0 citations
Preprint Aug 2026

Cross-Domain Generalization in Machine Unlearning via Label-Conditioned Energy Magnitude Regularization

This paper studies what happens to the rest of the model when a class is forgotten, using a label-conditioned energy-based model (EBM) that assigns per-class energies, making the effect directly observable.

Syed Ali Ahmed, Syed Bilal Ahsan, Muhammad Zaigham Zaheer National University of Computer et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.