This work proposes a unified compression framework, SaP, that jointly leverages basis sharing and coefficient pruning to reduce redundancies in both convolutional filters and their resulting feature maps and highlights the potential of combining complementary compression strategies to create highly efficient, deployable neural networks.
Abstract
Network compression is essential for deploying deep neural networks on resource-constrained devices, yet existing approaches often focus on either filter pruning or weight sharing in isolation. In this work, we propose a unified compression framework, SaP, that jointly leverages basis sharing and coefficient pruning to reduce redundancies in both convolutional filters and their resulting feature maps. By representing filters as linear combinations of a reduced set of shared basis filters and removing redundant coefficients, SaP achieves higher compression ratios while preserving discriminative information. This hybrid design mitigates the bottleneck problem inherent in sole basis-sharing methods, reduces computational complexity, and enables efficient fine-tuning in the coefficient subspace. We further provide a theoretical analysis showing that, under the orthonormal basis induced by the singular value decomposition, distances between filters are preserved in the coefficient space, justifying redundancy analysis and pruning in this low-dimensional representation. Extensive experiments on image classification, object detection, instance segmentation, and keypoint detection demonstrate that SaP consistently outperforms state-of-the-art baselines in terms of compression, accuracy retention, and practical efficiency. Ablation studies further validate the effectiveness of the joint approach and provide insights into hyperparameter selection and feature preservation. Our results highlight the potential of combining complementary compression strategies to create highly efficient, deployable neural networks.
Modern transformers pair impressive capabilities with substantial memory and compute demands. Low-rank weight factorization reduces both while keeping the matrices dense, and thus efficient on standard hardware. Existing methods, however, choose the subspace to remove from each weight matrix with local closed-form crit...
Massimo Bini, Anders Christensen, S. Alaniz et al.· 0 citations
Structured filter pruning is an effective approach for reducing the computational cost of deep neural networks while preserving efficient deployment on standard hardware. However, most existing pruning methods evaluate filters independently, overlooking the inter-layer spectral coupling between consecutive layers, whil...
Soumia Laroui, H. Seba, K. Amrouche et al.· Jordanian Journal of Compute...· 0 citations
A PyTorch-based, research-oriented framework is introduced that incorporates freezing and pruning mechanisms for designing and optimizing binarized neural networks and a novel pruning method that accounts for the relative importance of learned parameters across abstraction levels is proposed.
SVD-based pruning and quantization have recently emerged as a promising strategy for the ultra-efficient compression of large language models. In these methods, compression is performed in two stages: components are first truncated, and the remaining ones are subsequently quantized. Although this decoupled pipeline ben...
Deploying deep convolutional neural networks in resource-constrained edge environments necessitates aggressive model compression. While iterative block-level pruning paired with multi-stage Knowledge Distillation (KD) is a common strategy, traditional KD approaches rigidly enforce static, uniform loss weightings, leadi...
Chi-Dieu Bui, Chien-Quang Le· International Conference on...· 0 citations
Model compression is widely used to deploy large neural networks on resource-constrained edge devices. Among existing techniques, low-rank composition is theoretically grounded in approximation theory and provides a strong basis for preserving model performance after compression. However, in practice, even state-of-the...
Ya-Ping He, Hao Wu, Wei-Bo Liu et al.· Neural Networks· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.