This work treats the evolving SGD trajectory as a sequential experiment whose observations provide evidence about the unknown optimization error, and develops new recursive confidence-sequence techniques and a general time-uniform empirical Bernstein inequality for adapted processes with time-varying conditional means and predictable ranges that may grow without bound.
Abstract
Stochastic gradient descent (SGD) is typically analyzed at a deterministic horizon chosen before the algorithm is run, even though practical stopping decisions are made adaptively by inspecting the evolving trajectory. This mismatch creates a fundamental certification problem: fixed-time guarantees do not generally remain valid at data-dependent stopping times, while deterministic horizons derived from worst-case bounds can be highly conservative. We address this problem for strongly convex stochastic optimization by constructing fully observable, trajectory-adaptive upper confidence sequences for the squared distance of the last iterate to the optimizer and the suboptimality of a weighted average. These bounds hold simultaneously over time, attain the optimal $1/t$ decay rate up to iterated-logarithmic factors in the worst case, and adapt to the realized stochastic gradients, allowing SGD to stop as soon as a prescribed accuracy is certified without sacrificing statistical validity. Our approach treats the evolving SGD trajectory as a sequential experiment whose observations provide evidence about the unknown optimization error. To formalize this perspective, we develop new recursive confidence-sequence techniques and a general time-uniform empirical Bernstein inequality for adapted processes with time-varying conditional means and predictable ranges that may grow without bound. We further extend these confidence-sequence constructions to minibatch SGD, with the empirical Bernstein bounds exploiting the realized second-moment structure within each minibatch. Numerical experiments show that the resulting stopping rules can require several orders of magnitude fewer iterations than natural deterministic horizons.
The discrepancy principle is established as an adaptive regularization strategy for continuous-time SGD that is minimax-optimal up to a logarithmic factor, despite the persistent fluctuations induced by stochastic sampling.
Tim Jahn, Loucas Pillaud-Vivien, Adrien Schertzer· 0 citations
It is shown that, even for the covariance steering problem with a broad class of commonly used state and control safety constraints, the synthesized Markovian policy almost surely produces the same control actions as the history-dependent policy and therefore the same state trajectories, cost, and moments.
This paper provides the first finite-time convergence guarantees for this algorithm in this setting, for which it is proved that NPG converges sublinearly with a rate of $\mathcal{O}(H^{2}/t)$ after $t$ iterations, where $H$ is the horizon length.
Asha Barua, S. Khodadadian· arXiv.org· 0 citations
We study the quadratic tracking problem of a general stochastic target process with absolutely continuous controls, with and without terminal constraint. We derive explicit, non-asymptotic upper bounds in terms of a Besov-type modulus of the target. These bounds yield sharp explicit rates that specialize to the square-...
We study policy optimization for gain-scheduled linear quadratic regulation, where one schedule of gains, interpolated through fixed weighting functions, is optimized against a family of plants. The resulting cost can develop spurious local minima, and existing convergence certificates are either local or severely cons...
Shiva Shakeri, Péter Baranyi, M. Mesbahi· 0 citations
Expected survival time for a fixed deployed solution under isotropic Gaussian environmental dynamics is studied, derived from a rigorous lower bound and a computable multi-step upper bound, and an analytical characterization of deployment lifetime is provided.
Pavel Novoa-Hernández· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.