Aug 2026· Remote Sensing· 0 citations· 18 references
TL;DR
HSAR-DETR is proposed, a detection framework that jointly improves hierarchical feature representation, cross-scale refinement, and geometry-aware localization and experimental results on the VisDrone, RSOD, and TinyPerson datasets demonstrate improved detection performance.
Abstract
Small object detection in UAV remote sensing imagery plays a crucial role in applications such as infrastructure inspection, disaster assessment, and precision agriculture, where targets of interest frequently occupy fewer than 32×32 pixels under large ground sampling distance variation and complex cluttered backgrounds. Existing methods still face three main challenges in UAV small-object detection: fine-grained detail loss caused by repeated downsampling, feature inconsistency during cross-scale fusion, and unstable boundary regression in densely distributed aerial scenes. To address these issues, this paper proposes HSAR-DETR, a detection framework that jointly improves hierarchical feature representation, cross-scale refinement, and geometry-aware localization. Specifically, a Hierarchical Enhancement Network (HENet) is introduced to preserve shallow spatial details while strengthening deep semantic-context representation. A Dual-Stream Feature Refinement module (DSFR) is designed at the P4-to-P3 fusion stage, combining spatial-domain structural modeling with frequency-domain phase refinement to improve cross-scale feature consistency. A Coordinate-Guided Adaptive Convolution module (CGAC) is further deployed before the detection head, converting coordinate-guided offset magnitudes into modulation weights for adaptive feature recalibration and improved localization stability. In addition, a conventional high-resolution P2 detection branch is incorporated to enhance small-object representation. Experimental results on the VisDrone, RSOD, and TinyPerson datasets demonstrate improved detection performance. On the VisDrone validation set, HSAR-DETR achieves 50.8% mAP50 and 31.4% mAP50:95, outperforming the RT-DETR baseline by 4.2 and 3.0 percentage points, respectively.
Unmanned aerial vehicle (UAV) imagery is a core data source for remote sensing interpretation, intelligent transportation, urban monitoring, and disaster assessment, yet its large scale variation, dense object distributions, and complex backgrounds continue to challenge automated detection systems. Transformer-based de...
Shan Dan, Zan-Qi Qiu, Da-Di Cai et al.· Italian National Conference...· 0 citations
Small-object detection in unmanned aerial vehicle (UAV) remote sensing imagery is challenged by dense target distributions, substantial scale variation, complex ground backgrounds, and limited edge-computing resources. To address these challenges, we propose CDF-DETR, an end-to-end detector derived from the Real-Time D...
Unmanned aerial vehicle (UAV)-based remote sensing object detection faces three fundamental bottlenecks: (1) insufficient resolution diversity in single-scale detection heads, causing irreversible spatial detail loss for small targets; (2) semantic gap accumulation in multi-scale feature fusion due to content-agnostic...
Long Zhang· Journal of King Saud Univers...· 0 citations
Small-object detection in unmanned aerial vehicle (UAV) imagery remains challenging because target objects often occupy only a few pixels, exhibit weak feature responses, and are easily obscured by complex backgrounds. These aspects significantly limit the effectiveness of end-to-end detection systems. To overcome thes...
Object detection in unmanned aerial vehicle (UAV) imagery is severely challenged by extremely small object scales, cluttered backgrounds, and pronounced foreground–background imbalance, which jointly degrade the accuracy of general-purpose detectors. This paper presents HSF-Net, a small-object detection network built u...
Results verify that AARM effectively enhances the representation of small objects in complex remote sensing conditions of UAVs and achieves stable improvements in detection accuracy while maintaining high inference speed.
Yong Yang, Tian-Ci Wan, Meng-Lu Zhang· Signal, Image and Video Proc...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.