Skip to content

Gradient-Energy Guided Block-Wise Perturbations for Sharpness-Aware Minimization

Jul 2026 · arXiv.org · Vol abs/2607.18306 · 0 citations · 31 references
Computer Science

TL;DR

Gradient-Energy Adaptive Radius SAM (GEAR-SAM), which maintains an exponential moving average of squared block gradients as a lightweight, curvature-related sensitivity signal and allocates the fixed SAM budget through a closed-form constrained optimization, is proposed.

Abstract

Sharpness-Aware Minimization (SAM) improves generalization by minimizing the worst-case loss in a local parameter neighborhood. Standard SAM implicitly allocates its global perturbation budget across parameter blocks according to instantaneous minibatch gradient norms. Such an allocation can be noisy and may not reflect the sensitivity that blocks accumulate throughout training. We propose Gradient-Energy Adaptive Radius SAM (GEAR-SAM), which maintains an exponential moving average (EMA) of squared block gradients as a lightweight, curvature-related sensitivity signal and allocates the fixed SAM budget through a closed-form constrained optimization. GEAR-SAM preserves the global SAM radius, requires no Hessian-vector products or explicit Fisher estimation, and adds only scalar state beyond SAM. Experiments on image classification, transfer learning, noisy-label learning, and partition studies demonstrate improved generalization and robustness across architectures and tasks. More broadly, GEAR-SAM provides a dynamic view of sharpness-aware optimization: a fixed perturbation budget should be redistributed as the sensitivity of functional network blocks evolves during training.

View source

Similar papers

Conference Open access Aug 2026

EMASAM: a Computationally Efficient Sharpness-Aware Minimization via EMA-Guided Perturbations

Recent progress in optimization research has highlighted the sharpness of the loss landscape as a key factor in narrowing the generalization gap. Motivated by this insight, Sharpness-Aware Minimization (SAM) was proposed as a training strategy that enhances generalization. Despite the promising performance, SAM suffers...

Tanapat Ratchatorn, Masayuki Tanaka · 0 citations
Preprint Sep 2026

Sparsity-Adaptive Sharpness-Aware Minimization

Sparsity-Adaptive Sharpness-Aware Minimization (SA-SAM), which derives a sparsity-dependent SAM/ASAM perturbation radius by keeping the mean absolute perturbation approximately invariant as sparsity increases, is introduced.

Shiryu Ueno, Yoshikazu Hayashi, Kunihito Kato · 0 citations
Preprint Sep 2026

DA-Lion: Efficient Neural Video Representation via Direction-Aware Optimization

Direction-Aware Lion (DA-Lion), a task-driven optimizer tailored for NVR, introduces a direction-consistency criterion that switches between sign update and momentum update based on gradient--momentum alignment, and a learning-rate-aware magnitude modulation that stabilizes effective step sizes across training phases.

Qing-Yu Mao, Jia-Cong Chen, Shuai Liu et al. · 0 citations
#machine learning Preprint Sep 2026

Quantization-Aware Pre-Training with Constrained Empirical Weight Distribution

Quantization-Aware Pre-Training (QAPT) can increase the inference efficiency of DNNs, but a problematic behaviour known as rounding boundary weight oscillation can introduce detrimental noise into the training process and significantly reduce convergence speed. While existing methods can reduce this detrimental noise,...

Ning-Feng Yang, T. Aamodt · 0 citations
#machine learning Preprint Sep 2026

Explaining f-Divergence-Based Regularization via Local Curvature and Sharpness-Aware Minimization

Divergence-based regularization and Sharpness-Aware Minimization (SAM) are two prominent approaches for improving generalization in deep learning, both motivated by robustness to perturbations. However, their relationship has remained largely unexplored. Building on classical second-order expansions of $f$-divergences,...

Nour Jamoussi, Marios Kountouris · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.