Skip to content

Leveraging Extragradient for Effective Sharpness-Aware Minimization in Deep Learning

Jul 2026 · arXiv.org · Vol abs/2607.06151 · 0 citations · 57 references
Computer Science Mathematics

TL;DR

Theoretical analysis further confirms that EISAM tightens the generalization bound by steering parameters toward flatter minima with reduced curvature, establishing it as a robust, scalable, and broadly applicable optimization solution that advances both the theory and practice in deep learning.

Abstract

Generalization remains a pivotal challenge in deep learning, where traditional optimizers like Stochastic Gradient Descent (SGD) often converge to sharp minima, leading to overfitting and reduced performance on unseen data. Building on Sharpness-Aware Minimization (SAM), for seeking flat minima associated with improved generalization, we propose the Extragradient-Inspired Sharpness-Aware Minimization (EISAM), a novel optimizer that enhances generalization via the extragradient technique. EISAM uses a two-step update process: a prediction step investigating the geometry of the loss landscape and a perturbation step that refines updates with a base optimizer. This approach achieves better generalization performance than SAM. Crucially, EISAM reduces sensitivity to the perturbation radius, enhancing robustness, and simplifying the tuning across diverse settings. Extensive experiments on benchmark datasets demonstrate that EISAM consistently outperforms SGD, Adaptive Moment Estimation (Adam), and SAM in test accuracy and training efficiency across various architectures. Theoretical analysis further confirms that EISAM tightens the generalization bound by steering parameters toward flatter minima with reduced curvature. Accompanied by a thorough hyperparameter analysis, EISAM offers practical tuning guidance, establishing it as a robust, scalable, and broadly applicable optimization solution that advances both the theory and practice in deep learning.

View source

Similar papers

Open access Aug 2026

Quartz Optimizer: Robust Gradient Shaping and Bounded Adaptive Steps for Stable Deep Learning Training

Optimization plays a critical role in training deep neural networks, directly impacting convergence speed, model generalization, and stability. While existing methods such as stochastic gradient descent (SGD) and adaptive optimizers like Adam and AdamW have achieved significant success, they exhibit limitations in hand...

Ahmad Raza Khan, Sarab Almuhaideb · 0 citations
Conference Open access Aug 2026

EMASAM: a Computationally Efficient Sharpness-Aware Minimization via EMA-Guided Perturbations

Recent progress in optimization research has highlighted the sharpness of the loss landscape as a key factor in narrowing the generalization gap. Motivated by this insight, Sharpness-Aware Minimization (SAM) was proposed as a training strategy that enhances generalization. Despite the promising performance, SAM suffers...

Tanapat Ratchatorn, Masayuki Tanaka · 0 citations
Preprint Aug 2026

AOS: Adaptive Optimizer Switching via Training-State Signals for Faster Convergence and Better Generalization

AOS-R (Adaptive Optimizer Switching, Rule-Based), a lightweight controller that monitors six online gradient-space signals and switches among AdamW, SGD-M, and Lion as the optimization landscape evolves, and achieves best accuracy on 6 of 8 combinations with a mean +0.4 pp gain.

A. K. Pandey, Umang Chaturvedi, A. Rana et al. · 0 citations
Preprint Sep 2026

Sparsity-Adaptive Sharpness-Aware Minimization

Sparsity-Adaptive Sharpness-Aware Minimization (SA-SAM), which derives a sparsity-dependent SAM/ASAM perturbation radius by keeping the mean absolute perturbation approximately invariant as sparsity increases, is introduced.

Shiryu Ueno, Yoshikazu Hayashi, Kunihito Kato · 0 citations
Conference Aug 2026

Truncated SVD-BLS: Enforcing Flat Minima for Robust Broad Learning via Relative Flatness

Broad Learning System (BLS) achieves efficient training through random feature mapping and closed-form pseudoinverse solutions. However, the standard BLS suffers from limited generalization performance owing to its high sensitivity to feature perturbations, particularly when the feature matrix contains near-zero singul...

Yi-Xuan Gao, Liang-Ming Chen · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.