Efficient Model Pruning via Selective Layer-wise Distillation with Dynamic CKA-based Weighting
Abstract
Deploying deep convolutional neural networks in resource-constrained edge environments necessitates aggressive model compression. While iterative block-level pruning paired with multi-stage Knowledge Distillation (KD) is a common strategy, traditional KD approaches rigidly enforce static, uniform loss weightings, leading to representational collapse under severe depth pruning. We propose a novel Multi-stage Knowledge Distillation framework governed by a Dynamic Centered Kernel Alignment (CKA) weighting mechanism. By utilizing CKA to quantify the immediate post-pruning layer-wise representation similarity, our method topology-adaptively modulates distillation intensity, creating a representational "corridor of freedom". This empowers the "student" network to selectively filter crucial knowledge and optimize its feature space autonomously, preventing blind mimicry of the unpruned "teacher". Extensive experiments on closed-set image classification (CIFAR-10/100) and open-set face recognition finetuned on the CASIA-1k/3k and MS1MV2-1k datasets and evaluated on the LFW, CFP-FP, and AgeDB-30 benchmarks demonstrate remarkable scalability. Notably, our framework achieves a near-complete or full performance recovery on the evaluated benchmarks compared to the unpruned teacher, even under an extreme 74% depth reduction (20 blocks pruned). By slashing the recovery phase to just 60 epochs, our approach reduces standard fine-tuning time by 70% (from 200 to 60 epochs). With an Epoch Efficiency Score (Seff ) approximately 3.3× higher than traditional baselines, this work establishes a highly effective framework for rapid, high-performance model optimization in resource-constrained Edge AI environments.