Skip to content

Why Clipping Matters in AdaGrad? Toward a High-Probability Theory under Generalized Smoothness

Aug 2026 · 0 citations · 61 references
Computer Science Engineering

TL;DR

This work analyzes the original same-step coordinate-wise AdaGrad under generalized smoothness and heavy-tailed noise with bounded variance and shows that clipping is a structural stabilizer of the adaptive geometry rather than merely a robustness heuristic for AdaGrad under heavy-tailed noise.

Abstract

We analyze the original same-step coordinate-wise AdaGrad under generalized smoothness and heavy-tailed noise with bounded variance. In this setting, local curvature may grow sub-quadratically with the gradient norm, and stochastic gradients are assumed to have only bounded conditional second moments. We show that unclipped AdaGrad can become \emph{anisotropically miscalibrated}: under heavy-tailed noise, the adaptive denominator can learn the geometry of rare noise shocks rather than the local curvature of the objective, leading to a persistent directional distortion that blocks finite-horizon Euclidean progress. We then prove that clipping repairs this failure mode. Our main result is a finite-horizon high-probability guarantee for the original non-lagged AdaGrad update, yielding $\frac1T\sum_{t=0}^{T-1}\|\nabla f(x_t)\|^2=\mathcal{O}\left(\frac{d\big(\sqrt{\log T} + \log \frac{1}{\delta}\big)}{\sqrt{T}}\right),$ and hence $\widetilde{\mathcal O}(\varepsilon^{-2})$ complexity. This shows that, for AdaGrad under heavy-tailed noise, clipping is a structural stabilizer of the adaptive geometry rather than merely a robustness heuristic.

View source

Similar papers

#machine learning Preprint Sep 2026

Convergence of Stochastic Gradient Methods under Heavy-Tailed Noise and H\"{o}lder Smoothness

Classical convergence guarantees for stochastic gradient methods typically assume Lipschitz-smooth objectives and finite-variance gradient noise, both frequently violated in practice. In contrast, we study nonconvex stochastic optimization under the joint relaxation of these assumptions: objectives with $(L,s)$-H\"olde...

M. Zaman, Anirbit Mukherjee · 0 citations
#machine learning Preprint Oct 2026

Gradient-Free Sampling from Generative Models via Stochastic Bounded Extremum Seeking

We introduce a sampling approach for energy- and score-based generative models that requires no gradient evaluations of the model. Replacing the drift term that would normally contain the score $\nabla_\mathbf{x} \log p_\theta(\bf{x})$ with a high-frequency dithered cosine of the model's \textit{value}, $\sqrt{\alpha\o...

A. Scheinker · 0 citations
Preprint Sep 2026

Accelerated Stochastic Method under $(H_0,H_1)$-Smoothness and Heavy-Tailed Noise

We develop an accelerated stochastic method for convex $(H_0,H_1)$-smooth optimization under heavy-tailed noise. The unbiased oracle has a finite $p$-th noise moment, $1<p\le2$, with constant, gradient-dependent, and gap-dependent terms. Our method combines accelerated updates with clipping, projection, and phase resta...

A. Lobanov, D. Dvinskikh, A. Gasnikov · 0 citations
Preprint Sep 2026

Localize, Restart, Accelerate: Stochastic Optimization under Generalized Smoothness

A two-phase accelerated method that achieves, with high probability, an accelerated optimization contribution and smooth-subclass-optimal statistical dependence on accuracy, up to logarithmic and generalized-smoothness factors.

D. Dvinskikh, A. Gasnikov, A. Lobanov et al. · 1 citation
#machine learning Preprint Sep 2026

Ordinary Nonconvex SGD under Distance-Dependent Moments: Finite-Horizon Stationarity and Nagaev Bounds

Uniform noise-moment bounds exclude stochastic gradients whose variability increases with the iterate. We study ordinary, single-sample stochastic gradient descent for smooth, lower-bounded, possibly nonconvex objectives under distance-dependent conditional moments. Under second moments alone, a direct descent--displac...

Wei-Biao Wu · 0 citations

Related blog posts

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.