Skip to content
Preprint

Across the Loss Landscape with Progressive Growth

Aug 2026 · 0 citations · 40 references
Computer Science

TL;DR

Under standard local regularity conditions around non-degenerate minima, it is proved that local sublevel sets are well approximated by ellipsoids and that basin accessibility under frozen constraints can be characterized by an explicit effective curvature in the frozen directions.

Abstract

Deep neural networks generalize well despite their highly nonconvex, overparameterized loss landscapes, a phenomenon often associated with the geometry of the minima found by stochastic optimization. We study how incremental grow-and-optimize strategies bias training toward flatter regions by viewing growth as progressive constraint relaxation. Starting from a low-dimensional submodel, we iteratively expand the trainable parameters by unlocking nested random subspaces while freezing the orthogonal complement at the network initialization, re-optimizing after each expansion until the full architecture is reached. Under standard local regularity conditions around non-degenerate minima, we prove that local sublevel sets are well approximated by ellipsoids and that basin accessibility under frozen constraints can be characterized by an explicit effective curvature in the frozen directions. This leads to an explanation of the bias: progressive growth increases the relative weight of wide basins and suppresses sharp ones through a volume effect induced by the frozen constraints. We empirically validate these predictions in controlled toy landscapes and in a realistic ResNet/CIFAR-100 setting and confirm that although progressive subspace growth reliably produces flatter solutions, curvature reductions do not universally translate into improved test performance, highlighting subtleties in the flatness-generalization connection. The code is available at https://github.com/p0lcAi/Across-the-Loss-Landscape.

View source

Similar papers

Jul 2026

On the robustness of noisy solutions in non-convex neural networks

Using a finite energy message-passing algorithm, it is demonstrated numerically that thermal noise enables effective generalization in the regime of constraint densities where both recovering the teacher and finding a zero temperature solution are computationally hard.

Enrico M. Malatesta, A. Passalacqua, Riccardo Zecchina · 0 citations
Jul 2026

A Defense of the Quadratic Model

Due to the complexity of neural network loss landscapes, optimization theory is forced to rely on idealized models, and there is generally a tradeoff between how theoretically tractable the model is, and how accurately it describes the true optimization dynamics. In this work, we stress test the simplest possible model...

Alexandru Meterez, Pranav Ajit Nair, Depen Morwani et al. · 3 citations
Preprint Aug 2026

On the Principles Behind Neural Network Optimizers

This thesis develops a principled grounding for Adam and motivates new designs, and reveals new local structures in matrix-based nonconvex problems, and helps understand and improve recent NN optimizers, such as Muon.

Yu-Shun Zhang · 0 citations
Aug 2026

Optimal Initialization Scale for Neural Networks With Locally Quadratic Loss Landscapes: An SGD Dynamics Perspective.

Stochastic gradient descent (SGD), one of the most fundamental optimization algorithms in machine learning (ML), can be recast through a continuous-time approximation as a Fokker-Planck equation for Langevin dynamics, a viewpoint that has motivated many theoretical studies. Within this framework, we study the relations...

H. Horii, S. Has · 0 citations
Preprint Aug 2026

Branch Geometry and Finite-Radius Sensitivity of Hard-ReLU Training

This work characterize the intervening regime in which the perturbation radius is proportional to the GD step, and identifies sufficient response regimes: differentiating training is a choice of perturbation resolution as well as a choice of derivative.

Xiao-Yang Li, Run-Ni Zhou, Xin Yan · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.