Aug 2026· Machine Learning and Knowledge Extraction· 0 citations· 61 references
TL;DR
A new iterative pruning algorithm is proposed, Iterative Flow-Aware Pruning (IFAP), that leverages these measures to identify and eliminate non-essential parameters while preserving critical information pathways in deep neural networks.
Abstract
This paper presents a novel method for pruning deep neural networks based on the concept of flow, derived from the continuous modeling of signal propagation across layers. We derive flow functions for fully connected, convolutional, and self-attention architectures, and we propose a new iterative pruning algorithm, Iterative Flow-Aware Pruning (IFAP), that leverages these measures to identify and eliminate non-essential parameters while preserving critical information pathways. Extensive experiments across ten prominent architectures (including CNNs, vision transformers, and efficient mobile networks) on ten benchmark datasets demonstrate consistent accuracy–compression trade-offs: 81% of the evaluated configurations achieve a 60–81% reduction in computational cost relative to the corresponding baseline model. Furthermore, 97% of the evaluated configurations retain more than 98% of their baseline Top-1 accuracy. These results validate flow-based importance scoring as a robust and general-purpose foundation for model optimization.
This paper proposes a simple, novel and energy-efficient structured pruning methodology resulting in energy, memory footprint and parameter reduction and provides practical insights into the energy-accuracy tradeoff.
Diffa G. Pinto, Sunil B. Mane· Discover Computing· 0 citations
Deploying deep convolutional neural networks in resource-constrained edge environments necessitates aggressive model compression. While iterative block-level pruning paired with multi-stage Knowledge Distillation (KD) is a common strategy, traditional KD approaches rigidly enforce static, uniform loss weightings, leadi...
Chi-Dieu Bui, Chien-Quang Le· International Conference on...· 0 citations
This survey presents a comprehensive review of over 150 ViT variants and hybrid models, moving beyond application‐based categorizations to propose a multi‐dimensional taxonomy grounded in architectural evolution and attention design, and critically examines how different mechanisms influence scalability, computational...
Findings confirm that combining complementary compression strategies yields substantially better performance-efficiency trade-offs than any single technique applied in isolation.
Upma Sharma Archana· International Journal of Res...· 0 citations
This work proposes a unified compression framework, SaP, that jointly leverages basis sharing and coefficient pruning to reduce redundancies in both convolutional filters and their resulting feature maps and highlights the potential of combining complementary compression strategies to create highly efficient, deployabl...
A novel adaptive fusion framework that adaptively combines CNN and Transformer features through learnable gating, attention-based feature integration, and explainable-AI methods is developed, intended to improve both computational efficiency and model interpretability.
Komal Sharma, Monika Sainger· International journal of com...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.