This paper proposes a simple, novel and energy-efficient structured pruning methodology resulting in energy, memory footprint and parameter reduction and provides practical insights into the energy-accuracy tradeoff.
Abstract
Deep Convolutional Neural Networks (CNNs) have gained immense popularity over the past decade due to their exceptional performance, especially in imaging applications. However, the primary focus has been on improving model accuracy, often overlooking the significant environmental and computational costs associated with it. Pruning, a model compression technique, has been popularly used to reduce model size and computational complexity. Despite rapid growth of interest in this topic, research which comprehensively studies the energy consumption and targets energy efficiency using structured pruning is to date still missing. In this paper, we propose a simple, novel and energy-efficient structured pruning methodology resulting in energy, memory footprint and parameter reduction. Using the above methodology, an energy reduction of ranging from 6.49% to 58.58% is achieved, depending on the architecture and dataset. For VGG16 on CIFAR-10, the proposed methodology brings about 58% inference energy reduction, 89% reduction in memory footprint and 54% latency reduction for a negligible drop (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\sim 1\%$$\end{document}) in accuracy. The one-shot version of the energy-aware pruning algorithm achieves an impressive 74.64% reduction in training energy for VGG16 on CIFAR-10, leading to substantial additional energy savings. We also provide practical insights into the energy-accuracy tradeoff and encourage researchers to choose a suitable approach for pruning based on the required accuracy rather than best possible accuracy for a particular application to ensure sustainable and Green AI.
A new iterative pruning algorithm is proposed, Iterative Flow-Aware Pruning (IFAP), that leverages these measures to identify and eliminate non-essential parameters while preserving critical information pathways in deep neural networks.
A. Samarin, Artem A. Nazarenko, E. Kotenko et al.· Machine Learning and Knowled...· 0 citations
Findings confirm that combining complementary compression strategies yields substantially better performance-efficiency trade-offs than any single technique applied in isolation.
Upma Sharma Archana· International Journal of Res...· 0 citations
Deep neural networks excel in processing vision tasks, but their high computational complexity prevents them from being deployed directly on edge devices. In this paper, the ResNet18 model is optimized on cifar-10 data set by structured pruning and quantization-aware training (QAT). This study compared the model accura...
Edge computing scenarios put forward strict requirements on inference delay and computational power consumption of target detectors, and the existing lightweight models are difficult to achieve the best balance between accuracy and speed. In this paper, a hybrid architecture edge object detector based on structured pru...
Rui-Sheng Zhang· Journal of Discovery Core· 1 citation
It is demonstrated that hyperparameter optimization dynamics depend heavily on dataset complexity, where computational efficiency is the primary differentiator for simpler classification tasks, but optimization architecture selection becomes critical for navigating challenging medical imaging applications.
Sarab Almuhaideb, Ahmad Raza Khan· Applied Sciences· 0 citations