LASAD-YOLO: Localization and Adaptive Spatial Attention Distillation for Dense Object Detection
Abstract
This article introduces a novel object detection framework that integrates localization and adaptive spatial attention distillation techniques. While effective, prior knowledge distillation (KD) methods face a critical challenge in harmonizing feature-based and logit-based philosophies, often providing either entangled knowledge or incomplete guidance for modern detectors. This research aims to address this gap by developing a comprehensive and unified distillation approach. By concurrently training teacher and student networks, the proposed method significantly improves the student model’s detection accuracy without adding extra computational demands during inference, as evidenced by identical parameter counts and floating-point operation per seconds (FLOPs). The research aims to develop a comprehensive and unified KD approach for visual object detection, leveraging teacher-guided attention to effectively balance feature-based and logit-based distillation methods. Knowledge transfer is optimized by distinctly separating the processes into localization distillation (LD) and classification distillation derived from feature maps. The robustness and strong generalization capabilities of this method are demonstrated through extensive experiments on the MS-COCO 2017 dataset, with comparisons against existing approaches. The results consistently show performance improvements across different model scales, achieving up to a 0.9% enhancement over baseline models. The best-performing model records an average precision (AP) of 53.1% on the MS-COCO dataset.