LAI-YOLO: a lightweight attention-integrated framework for small object detection in UAV imagery
Abstract
Identifying miniature targets within drone-captured visuals continues to be a demanding task, primarily driven by restricted image clarity, intricate environmental contexts, and the scarce computational power available on flight platforms. In response to such bottlenecks, we introduce a novel architecture termed LAI-YOLO. By integrating attention mechanisms into a compact design, this model successfully strikes an optimal equilibrium between computational swiftness and recognition precision. The key improvements are summarized as follows: 1) In the backbone, a Cross-stage partial Shiftwise Partitioning Dynamic Patch-aware attention (C_SPDP) block is introduced to enhance the detection accuracy. 2) In the neck, a Hierarchical Feature Pyramid Network (HFPN) is constructed to reduce model parameters. 3) For the detection head, we introduce a Lightweight Scale-aware Convolutional Detection with Localization Quality Estimation (LSCD_LQE) structure to reduce computational overhead and enhance localization consistency. 4) Moreover, a new loss function, Wise-Inner-ShapeIoU (WISIoU), is developed to refine bounding-box regression through shape consistency constraints and adaptive gradient scaling. The proposed LAI-YOLO achieves 34.8% mAP50 on the VisDrone2019 dataset with only 2.38M parameters. Furthermore, its generalization capability is validated on the UAVDT, AI-TOD, and CODrone datasets, demonstrating the robustness and effectiveness of our method.