Skip to content

RELTA-SGLD: Relative-Growth Localized Taming for Nonconvex Stochastic-Gradient Langevin Learning

Jul 2026 · arXiv.org · Vol abs/2607.19544 · 0 citations · 26 references
Mathematics Computer Science

TL;DR

RELTA-SGLD is introduced, a taming scheme that stabilizes superlinear stochastic-gradient updates while reducing unnecessary suppression of the original learning drift and maintains nearly untamed learning dynamics.

Abstract

We introduce RELTA-SGLD, a taming scheme that stabilizes superlinear stochastic-gradient updates while reducing unnecessary suppression of the original learning drift. A threshold determines where the taming turns on, while a relative-growth principle derived from the one-step Lyapunov stability condition determines the required taming strength. Together, they produce a lighter $\lambda$-scale denominator and preserve a nonvanishing far-tail return. As a consequence, we prove polynomial moment stability and first-order stationary accuracy in both $W_1$ and $W_2$ for nonconvex SGLD with superlinearly growing stochastic-gradient oracles, improving the corresponding half-order and quarter-order bounds for comparable stochastic-gradient tamed schemes. On Fashion-MNIST under active stabilization pressure, RELTA improves the mean learning metrics over both untamed SGLD and TUSLA and remains competitive with a tuned AdamW reference. In an ordinary-training regime, its lighter localized denominator reduces unnecessary perturbation of the original update and maintains nearly untamed learning dynamics.

View source

Similar papers

Preprint Aug 2026

Adaptive Bregman Proximal Stochastic Gradient with a Stabilized Barzilai--Borwein Step Size

Ada-BPSG is introduced, a line-search-free BPSG method that couples the SAGA gradient table with a stabilized Barzilai--Borwein (BB) candidate and yields a direct analytical chain from relative smoothness and component-wise variance control to convergence in finite-dimensional normed spaces.

Chen-Han Jin, Sheng-Ze Xu, Bing-Hui Xie et al. · 0 citations
#machine learning Preprint Sep 2026

Exact information accounting for SGD methods

As an alternative to the standard geometric analyses, we give an exact, information-theoretic analysis of stochastic gradient descent (SGD) and its variants. We show that a preconditioned SGD step is the posterior-mean update of a Gaussian Bayes model, and that its one-step regret splits into an intrinsic-time cost and...

Akshay Balsubramani · 0 citations
#machine learning Preprint Sep 2026

High-Probability Convergence of SGD via Batched Updates

Stochastic gradient descent (SGD) is the primary workhorse for large-scale optimization. While the average behavior of its iterates, typically characterized by mean-squared error bounds, is well-understood, obtaining high-probability guarantees for the last iterate remains challenging. Prior approaches to this problem...

Feng Zhu, Robert W. Heath, Aritra Mitra · 0 citations
#artificial intelligence Preprint Sep 2026

Adam under Generalized Smoothness with Second-Moment-Type Stochastic Gradients

Adam is widely observed to remain stable even when the objective deviates significantly from global smoothness. Under the generalized smoothness framework, however, existing analyses rely on strong tail assumptions on the stochastic gradients, such as almost-sure boundedness or sub-Gaussianity. Whether Adam converges o...

Rui-Nan Jin, Di-Fei Cheng, Ling Chen et al. · 0 citations
#artificial intelligence Preprint Sep 2026

G$^2$PTQ: Improving LLM Post-Training Quantization with Generalized Gradient Compensation

G$^2$PTQ is presented, a unified PTQ framework with Generalized Gradient Compensation that integrates both first- and second-order information under a globally supervised, block-wise optimization objective and enables better alignment with the full-precision model, outperforming state-of-the-art baselines.

Rui-Kang Liu, Hao-Li Bai, Yuxuan Sun et al. · 0 citations
Preprint Sep 2026

Accelerated Stochastic Method under $(H_0,H_1)$-Smoothness and Heavy-Tailed Noise

We develop an accelerated stochastic method for convex $(H_0,H_1)$-smooth optimization under heavy-tailed noise. The unbiased oracle has a finite $p$-th noise moment, $1<p\le2$, with constant, gradient-dependent, and gap-dependent terms. Our method combines accelerated updates with clipping, projection, and phase resta...

A. Lobanov, D. Dvinskikh, A. Gasnikov · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.