Skip to content

Basis sharing and coefficient pruning for efficient neural network compression.

Sep 2026 · Neural Networks · Vol 205 Pt C, pp. 109630 · 0 citations · 95 references
Medicine

TL;DR

This work proposes a unified compression framework, SaP, that jointly leverages basis sharing and coefficient pruning to reduce redundancies in both convolutional filters and their resulting feature maps and highlights the potential of combining complementary compression strategies to create highly efficient, deployable neural networks.

Abstract

Network compression is essential for deploying deep neural networks on resource-constrained devices, yet existing approaches often focus on either filter pruning or weight sharing in isolation. In this work, we propose a unified compression framework, SaP, that jointly leverages basis sharing and coefficient pruning to reduce redundancies in both convolutional filters and their resulting feature maps. By representing filters as linear combinations of a reduced set of shared basis filters and removing redundant coefficients, SaP achieves higher compression ratios while preserving discriminative information. This hybrid design mitigates the bottleneck problem inherent in sole basis-sharing methods, reduces computational complexity, and enables efficient fine-tuning in the coefficient subspace. We further provide a theoretical analysis showing that, under the orthonormal basis induced by the singular value decomposition, distances between filters are preserved in the coefficient space, justifying redundancy analysis and pruning in this low-dimensional representation. Extensive experiments on image classification, object detection, instance segmentation, and keypoint detection demonstrate that SaP consistently outperforms state-of-the-art baselines in terms of compression, accuracy retention, and practical efficiency. Ablation studies further validate the effectiveness of the joint approach and provide insights into hyperparameter selection and feature preservation. Our results highlight the potential of combining complementary compression strategies to create highly efficient, deployable neural networks.

View source

Similar papers

#machine learning Preprint Sep 2026

Learning Functional Subspaces for Neural Network Compression

Modern transformers pair impressive capabilities with substantial memory and compute demands. Low-rank weight factorization reduces both while keeping the matrices dense, and thus efficient on standard hardware. Existing methods, however, choose the subspace to remove from each weight matrix with local closed-form crit...

Massimo Bini, Anders Christensen, S. Alaniz et al. · 0 citations
Open access 2026

Filter Spectral Efficiency (FSE): Structured Network Pruning via Inter-Layer Spectral Coupling

Structured filter pruning is an effective approach for reducing the computational cost of deep neural networks while preserving efficient deployment on standard hardware. However, most existing pruning methods evaluate filters independently, overlooking the inter-layer spectral coupling between consecutive layers, whil...

Soumia Laroui, H. Seba, K. Amrouche et al. · 0 citations
Preprint Aug 2026

Pruning Binarized Neural Networks: A Dedicated Framework and Globally Weighted Algorithms

A PyTorch-based, research-oriented framework is introduced that incorporates freezing and pruning mechanisms for designing and optimizing binarized neural networks and a novel pruning method that accounts for the relative importance of learned parameters across abstraction levels is proposed.

Roan Rubiales, Jean-Pierre David · 0 citations
#artificial intelligence Preprint Oct 2026

Differentiable Bit-Widths: Co-optimizing Pruning and Quantization via SVD for Ultra-Efficient LLM Compression

SVD-based pruning and quantization have recently emerged as a promising strategy for the ultra-efficient compression of large language models. In these methods, compression is performed in two stages: components are first truncated, and the remaining ones are subsequently quantized. Although this decoupled pipeline ben...

Hankyul Kang, Jongbin Ryu · 0 citations
Conference Aug 2026

Efficient Model Pruning via Selective Layer-wise Distillation with Dynamic CKA-based Weighting

Deploying deep convolutional neural networks in resource-constrained edge environments necessitates aggressive model compression. While iterative block-level pruning paired with multi-stage Knowledge Distillation (KD) is a common strategy, traditional KD approaches rigidly enforce static, uniform loss weightings, leadi...

Chi-Dieu Bui, Chien-Quang Le · 0 citations
Open access Sep 2026

A novel approach to lossless convolutional neural network compression via progressive knowledge distillation-incorporated low-rank compression.

Model compression is widely used to deploy large neural networks on resource-constrained edge devices. Among existing techniques, low-rank composition is theoretically grounded in approximation theory and provides a strong basis for preserving model performance after compression. However, in practice, even state-of-the...

Ya-Ping He, Hao Wu, Wei-Bo Liu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.