SPO-YOLO: Lightweight Small-Object Detection for UAV Aerial Imagery with Federated Learning
Abstract
Detecting small objects in unmanned aerial vehicle (UAV) aerial imagery remains challenging due to tiny target scales, cluttered backgrounds, and strict onboard resource constraints. We propose SPO-YOLO, a lightweight small-object-oriented detector built upon YOLOv11n. SPO-YOLO introduces (i) a P2 high-resolution detection head to preserve spatial details for tiny targets, (ii) spatial-to-depth convolution (SPDConv) in early backbone stages to reduce information loss during downsampling, and (iii) a lightweight small object detection attention (SODA) attention module that enhances small-object responses on high-resolution features while suppressing background redundancy. To enable privacy-preserving multi-UAV collaboration, we further integrate SPO-YOLO into a federated learning pipeline for cross-scene training without sharing raw data. On VisDrone2019-DET, SPO-YOLO improves mAP 50 by 7.0% over YOLOv11n while maintaining a compact model footprint. Cross-dataset evaluation on UAVDT shows a 5.1% gain in mAP 50 , indicating improved generalization. Under federated training, FL-SPOYOLO exceeds the Federated Learning baseline by 6.4% in mAP 50 , demonstrating the effectiveness of the proposed model.