Skip to content
Preprint

Domain-Aware Pruning: Sparsity and Domain Generalization via Regularized Probabilistic Masking

Aug 2026 · 0 citations · 40 references
Computer Science

TL;DR

Beyond efficiency and generalization, DAP natively provides increased robustness to adversarial perturbations and yields highly interpretable models, where the retained weights reliably encapsulate the most domain-invariant and task-critical representations.

Abstract

Domain generalization (DG) and neural network pruning are conventionally treated as distinct objectives, targeting out-of-distribution (OOD) robustness and model efficiency, respectively. In this work, we bridge this gap by introducing Domain-Aware Pruning (DAP), a framework that leverages network sparsity as a mechanism to implicitly enhance generalization to unseen domains. Diverging from standard binary mask optimization, DAP learns a continuous parameter retention probability $p \in [0, 1]$, framing network compression as a continuous probabilistic masking problem. By introducing a regularization objective that actively penalizes the retention of domain-sensitive weights during the mask training, DAP identifies a domain-invariant subnetwork. Empirical results across five DG benchmark datasets demonstrate that DAP achieves significant sparsity while consistently matching or exceeding the OOD performance of its dense counterparts. Crucially, DAP is an algorithm-agnostic framework that integrates seamlessly with existing DG pipelines without necessitating post-hoc fine-tuning. Beyond efficiency and generalization, we show that DAP natively provides increased robustness to adversarial perturbations and yields highly interpretable models, where the retained weights reliably encapsulate the most domain-invariant and task-critical representations.

View source

Similar papers

Preprint Sep 2026

Sparsity-Adaptive Sharpness-Aware Minimization

Sparsity-Adaptive Sharpness-Aware Minimization (SA-SAM), which derives a sparsity-dependent SAM/ASAM perturbation radius by keeping the mean absolute perturbation approximately invariant as sparsity increases, is introduced.

Shiryu Ueno, Yoshikazu Hayashi, Kunihito Kato · 0 citations
Book Open access Aug 2026

Sparse Additive Models for Domain Generalization

This work incorporates an additive structure into the DG framework and employs ℓq,1 -norm regularization to induce sparsity, thereby enabling structured feature selection and enhancing interpretability and presents two distinct realizations: an additive kernel-based formulation and a neural additive model-based approac...

Jia-Yi Wang, Han Li · 0 citations
Preprint Aug 2026

Reservoir of Importance: Learning Semi-Structured Sparsity with Differentiable Subset Sampling

Semi-structured $N$:$M$ sparsity has emerged as a practical direction for accelerating large language models (LLMs). However, existing learnable-mask approaches incur substantial parameter and memory overhead, limiting their scalability to large models and aggressive sparsity regimes. In this work, we revisit semi-stru...

H. Dinh, Xuan Duy Ta, K. Thân et al. · 0 citations
Conference Open access Sep 2026

Diverge to Converge: Mutual Heterogeneous Learning for Robust Pruning

Mutual Heterogeneous Learning (MHL) is proposed, a framework enabling robust pruning via single-model inference that significantly outperforms single-model baselines in both adversarial robustness and corruption robustness, while maintaining competitive clean accuracy.

Jin-Hui Yu, Zikai Zhang, Khaled A. Harras et al. · 0 citations
Aug 2026

Random Sparse Networks Training with Sharpness-Aware Regularization.

Over-parameterization is critical for optimizing neural networks, whereas training sparse networks directly often fails to achieve satisfactory performance. However, the Lottery Ticket Hypothesis (LTH) demonstrates that a randomly initialized dense model has a sparse subnetwork that can be identified through iterative...

Yue Bai, Mingyuan Zhang, Huan Wang et al. · 0 citations
#machine learning Preprint Aug 2026

Correlation-Aware Structured Pruning for Large Language Models

Structured pruning is a promising approach for reducing the substantial inference costs of Large Language Models (LLMs) while maintaining hardware efficiency. Many existing methods assess the importance of prunable units (e.g., channels or heads) in isolation, implicitly assuming that pruning errors are additive. This...

Si-Cheng Xu, Hao Shi, Wei Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.