Sparsity-Adaptive Sharpness-Aware Minimization (SA-SAM), which derives a sparsity-dependent SAM/ASAM perturbation radius by keeping the mean absolute perturbation approximately invariant as sparsity increases, is introduced.
Abstract
Deploying deep neural networks in real-world settings requires models that are both compact and robust to common corruptions. However, at deployment-relevant high sparsity, standard pruning pipelines often degrade corruption robustness, and existing sharpness-aware training/pruning approaches provide limited robustness gains. We address this issue by introducing Sparsity-Adaptive Sharpness-Aware Minimization (SA-SAM), which derives a sparsity-dependent SAM/ASAM perturbation radius by keeping the mean absolute perturbation (an $\ell_1$-based proxy) approximately invariant as sparsity increases. As a simple complementary option, we evaluate Magnitude-Weighted Hessian (MWH), derived from a second-order removal-path analysis, yielding an importance proportional to $\mathrm{Diag}(F)_i\,|w_i|$, where $\mathrm{Diag}(F)$ is the diagonal empirical Fisher used as a curvature proxy in our implementation. Across CIFAR-10-C, CIFAR-100-C, and ImageNet-100-C, our approach achieved stronger corruption robustness than the considered pruning baselines at 80--90\% sparsity, while preserving clean accuracy. We additionally quantify the robustness--throughput trade-off by reporting measured inference throughput under sparse execution at deployment-relevant sparsity levels.
Over-parameterization is critical for optimizing neural networks, whereas training sparse networks directly often fails to achieve satisfactory performance. However, the Lottery Ticket Hypothesis (LTH) demonstrates that a randomly initialized dense model has a sparse subnetwork that can be identified through iterative...
Yue Bai, Mingyuan Zhang, Huan Wang et al.· IEEE Transactions on Pattern...· 0 citations
Beyond efficiency and generalization, DAP natively provides increased robustness to adversarial perturbations and yields highly interpretable models, where the retained weights reliably encapsulate the most domain-invariant and task-critical representations.
A systematic study of how pruning affects SAE behavior is presented and theoretically shows that, for a fixed SAE, its impact is governed by perturbation energy, a covariance-weighted norm.
Suchit Gupte, Xue-Ru Zhang, M. Khalili· 0 citations
Mutual Heterogeneous Learning (MHL) is proposed, a framework enabling robust pruning via single-model inference that significantly outperforms single-model baselines in both adversarial robustness and corruption robustness, while maintaining competitive clean accuracy.
Jin-Hui Yu, Zikai Zhang, Khaled A. Harras et al.· Proceedings of the Thirty-Fi...· 0 citations
Semi-structured $N$:$M$ sparsity has emerged as a practical direction for accelerating large language models (LLMs). However, existing learnable-mask approaches incur substantial parameter and memory overhead, limiting their scalability to large models and aggressive sparsity regimes. In this work, we revisit semi-stru...
H. Dinh, Xuan Duy Ta, K. Thân et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.