Skip to content
Preprint

Sparsity-Adaptive Sharpness-Aware Minimization

Sep 2026 · 0 citations · 20 references
Computer Science

TL;DR

Sparsity-Adaptive Sharpness-Aware Minimization (SA-SAM), which derives a sparsity-dependent SAM/ASAM perturbation radius by keeping the mean absolute perturbation approximately invariant as sparsity increases, is introduced.

Abstract

Deploying deep neural networks in real-world settings requires models that are both compact and robust to common corruptions. However, at deployment-relevant high sparsity, standard pruning pipelines often degrade corruption robustness, and existing sharpness-aware training/pruning approaches provide limited robustness gains. We address this issue by introducing Sparsity-Adaptive Sharpness-Aware Minimization (SA-SAM), which derives a sparsity-dependent SAM/ASAM perturbation radius by keeping the mean absolute perturbation (an $\ell_1$-based proxy) approximately invariant as sparsity increases. As a simple complementary option, we evaluate Magnitude-Weighted Hessian (MWH), derived from a second-order removal-path analysis, yielding an importance proportional to $\mathrm{Diag}(F)_i\,|w_i|$, where $\mathrm{Diag}(F)$ is the diagonal empirical Fisher used as a curvature proxy in our implementation. Across CIFAR-10-C, CIFAR-100-C, and ImageNet-100-C, our approach achieved stronger corruption robustness than the considered pruning baselines at 80--90\% sparsity, while preserving clean accuracy. We additionally quantify the robustness--throughput trade-off by reporting measured inference throughput under sparse execution at deployment-relevant sparsity levels.

View source

Similar papers

Aug 2026

Random Sparse Networks Training with Sharpness-Aware Regularization.

Over-parameterization is critical for optimizing neural networks, whereas training sparse networks directly often fails to achieve satisfactory performance. However, the Lottery Ticket Hypothesis (LTH) demonstrates that a randomly initialized dense model has a sparse subnetwork that can be identified through iterative...

Yue Bai, Mingyuan Zhang, Huan Wang et al. · 0 citations
Preprint Aug 2026

Domain-Aware Pruning: Sparsity and Domain Generalization via Regularized Probabilistic Masking

Beyond efficiency and generalization, DAP natively provides increased robustness to adversarial perturbations and yields highly interpretable models, where the retained weights reliably encapsulate the most domain-invariant and task-critical representations.

Parham Sazdar, Mostafa Tavassolipour, Reshad Hosseini · 0 citations
Conference Open access Sep 2026

Diverge to Converge: Mutual Heterogeneous Learning for Robust Pruning

Mutual Heterogeneous Learning (MHL) is proposed, a framework enabling robust pruning via single-model inference that significantly outperforms single-model baselines in both adversarial robustness and corruption robustness, while maintaining competitive clean accuracy.

Jin-Hui Yu, Zikai Zhang, Khaled A. Harras et al. · 0 citations
Preprint Aug 2026

Reservoir of Importance: Learning Semi-Structured Sparsity with Differentiable Subset Sampling

Semi-structured $N$:$M$ sparsity has emerged as a practical direction for accelerating large language models (LLMs). However, existing learnable-mask approaches incur substantial parameter and memory overhead, limiting their scalability to large models and aggressive sparsity regimes. In this work, we revisit semi-stru...

H. Dinh, Xuan Duy Ta, K. Thân et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.