Skip to content
Preprint

Pruning Binarized Neural Networks: A Dedicated Framework and Globally Weighted Algorithms

Aug 2026 · 0 citations · 17 references
Computer Science

TL;DR

A PyTorch-based, research-oriented framework is introduced that incorporates freezing and pruning mechanisms for designing and optimizing binarized neural networks and a novel pruning method that accounts for the relative importance of learned parameters across abstraction levels is proposed.

Abstract

Extreme compression of deep neural networks, up to full binarization, dramatically reduces memory footprint and arithmetic complexity, facilitating deployment on constrained edge hardware with field-programmable gate arrays (FPGAs) and microcontrollers. Although combining binarization with pruning promises additional efficiency gains, existing pruning strategies are ill-suited to binarized representations and rarely translate into meaningful hardware savings. We introduce a PyTorch-based, research-oriented framework that incorporates freezing and pruning mechanisms for designing and optimizing binarized neural networks. The framework enables rapid and reproducible evaluation of state-of-the-art approaches and the fast prototyping of new ones. Leveraging this framework, we propose a novel pruning method that accounts for the relative importance of learned parameters across abstraction levels. Such a global weighting mechanism consistently achieves a superior trade-off between model accuracy and pruning rate, achieving a 70% pruning rate on VGG11 with constant accuracy, while state-of-the-art results reach only 41% in the binarized setting.

View source

Similar papers

Open access 2023

Accelerating Neural Networks with Model Compression Techniques

Experimental results demonstrate that effective compression significantly reduces model size and computational cost with minimal performance loss, highlighting the importance of compression-aware design and concluding as a valuable reference for building efficient and scalable AI systems.

Daniel Rodríguez · 0 citations
Sep 2026

Basis sharing and coefficient pruning for efficient neural network compression.

This work proposes a unified compression framework, SaP, that jointly leverages basis sharing and coefficient pruning to reduce redundancies in both convolutional filters and their resulting feature maps and highlights the potential of combining complementary compression strategies to create highly efficient, deployabl...

Van Tien Pham · 0 citations
Open access Jul 2026

Design of Resource-Efficient AI Models through Parameter Reduction and Accuracy-Aware Compression

The proposed hybrid pipeline includes structured pruning, INT8 quantization and task-specific knowledge distillation, which is benchmarked against standalone methods and reinforces the idea of upper bound projection based approach for accuracy-oriented, multi-level compression.

Krishna Kumar Tiwari, Komal Tahiliani, Uma Shankar Birthare et al. · 0 citations
Preprint Aug 2026

APQF: Agentic Profiling-Guided Structured Pruning and Mixed-Precision Quantization with Adaptive Fine-Tuning

Modern deep neural networks achieve strong performance, but their scale makes them costly and slow, especially on resource-constrained edge devices. Pruning and quantization address this, but rely on manual, expert choices and on algorithms that are hard to apply across architectures. Uniform settings also ignore how d...

Sadegh Jafari, Mohiuddin Bilwal, Fansen Zhou et al. · 0 citations
Sep 2026

Adaptive multi-bit progressive quantization for stable training of binary neural networks

This novel progressive quantization framework combines multi-bit assistive teacher models with self-knowledge distillation to stabilize BNN training and integrates matching structured pruning with an asymmetric Binary Weight Network scaling factor, thereby reducing quantization errors while maintaining hardware efficie...

Jie Xu, Wonjun Hwang, Hyunsouk Cho · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.