Jul 2026
Effects of width-dependent model hyperparameters and ℓ2-regularization on the loss landscape of two-layer ReLU networks
Interestingly, the experiments reveal that using AdamW as an optimizer prevents the collapse of the learned parameters, whereas using SGD does not, which may help explain the success of AdamW in deep learning training.
Haruka Eshima, Makoto Yamada
· arXiv.org · 0 citations