Skip to content
Open access

Flow-Guided Neural Pruning: Signal-Flow Framework for Multi-Architecture Model Compression

Aug 2026 · Machine Learning and Knowledge Extraction · 0 citations · 61 references

TL;DR

A new iterative pruning algorithm is proposed, Iterative Flow-Aware Pruning (IFAP), that leverages these measures to identify and eliminate non-essential parameters while preserving critical information pathways in deep neural networks.

Abstract

This paper presents a novel method for pruning deep neural networks based on the concept of flow, derived from the continuous modeling of signal propagation across layers. We derive flow functions for fully connected, convolutional, and self-attention architectures, and we propose a new iterative pruning algorithm, Iterative Flow-Aware Pruning (IFAP), that leverages these measures to identify and eliminate non-essential parameters while preserving critical information pathways. Extensive experiments across ten prominent architectures (including CNNs, vision transformers, and efficient mobile networks) on ten benchmark datasets demonstrate consistent accuracy–compression trade-offs: 81% of the evaluated configurations achieve a 60–81% reduction in computational cost relative to the corresponding baseline model. Furthermore, 97% of the evaluated configurations retain more than 98% of their baseline Top-1 accuracy. These results validate flow-based importance scoring as a robust and general-purpose foundation for model optimization.

Read PDF

Similar papers

Open access Sep 2026

Towards energy-efficient CNNs using structured pruning

This paper proposes a simple, novel and energy-efficient structured pruning methodology resulting in energy, memory footprint and parameter reduction and provides practical insights into the energy-accuracy tradeoff.

Diffa G. Pinto, Sunil B. Mane · 0 citations
Conference Aug 2026

Efficient Model Pruning via Selective Layer-wise Distillation with Dynamic CKA-based Weighting

Deploying deep convolutional neural networks in resource-constrained edge environments necessitates aggressive model compression. While iterative block-level pruning paired with multi-stage Knowledge Distillation (KD) is a common strategy, traditional KD approaches rigidly enforce static, uniform loss weightings, leadi...

Chi-Dieu Bui, Chien-Quang Le · 0 citations
Review Sep 2026

The Evolution of Vision Transformers: A Multi‐Dimensional Analysis of Architectural Innovation and Application Domains

This survey presents a comprehensive review of over 150 ViT variants and hybrid models, moving beyond application‐based categorizations to propose a multi‐dimensional taxonomy grounded in architectural evolution and attention design, and critically examines how different mechanisms influence scalability, computational...

Shobhit Tyagi, Piyush Rawat, Divakar Yadav et al. · 0 citations
Sep 2026

Basis sharing and coefficient pruning for efficient neural network compression.

This work proposes a unified compression framework, SaP, that jointly leverages basis sharing and coefficient pruning to reduce redundancies in both convolutional filters and their resulting feature maps and highlights the potential of combining complementary compression strategies to create highly efficient, deployabl...

Van Tien Pham · 0 citations
Open access Aug 2026

Adaptive Feature Integration in CNN–Transformer Networks for Efficient and Interpretable Visual Classification

A novel adaptive fusion framework that adaptively combines CNN and Transformer features through learnable gating, attention-based feature integration, and explainable-AI methods is developed, intended to improve both computational efficiency and model interpretability.

Komal Sharma, Monika Sainger · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.