Efficient Compression of ResNet18 Model on CIFAR-10 Dataset
Abstract
Deep neural networks excel in processing vision tasks, but their high computational complexity prevents them from being deployed directly on edge devices. In this paper, the ResNet18 model is optimized on cifar-10 data set by structured pruning and quantization-aware training (QAT). This study compared the model accuracy, channel sparsity, theoretical model size, and other data under different pruning rates and quantitative accuracy combinations. The study found that the 50% pruning with the INT8 quantization strategy achieved the highest accuracy of 93.89%, which was 1.78% higher than the baseline model. The 80% pruning with the INT8 quantization strategy is most suitable for edge device deployment. Compared with the baseline model, the accuracy is only reduced by 1.15%, while the channel sparsity is 72.87%, and the model size is reduced by 75%. The study also found that the critical value of ResNet18 pruning was about 80%. This paper verifies the effectiveness of the structured pruning with the QAT optimization method and finds the optimization strategy corresponding to the model with the highest accuracy and the smallest volume, which provides a feasible scheme for the edge deployment of the ResNet18 model.