2026· IEEE Transactions on Instrumentation and Measurement· Vol 75, pp. 2515618-2515618· 0 citations· 55 references
Abstract
Reliable vision-based sensing of tiny-airborne objects is important for airspace surveillance and aerial monitoring. In practical scenarios, airborne objects are typically captured at long distances, occupying only a few pixels and exhibiting low signal-to-clutter ratios (SCRs) against complex backgrounds. These factors frequently cause missed detections and false alarms, thereby degrading reliable target detection and overall sensing reliability. Existing approaches seek to address these challenges by exploiting temporal cues across frames to stabilize weak object responses. However, they rarely impose explicit reliability-oriented verification on such cues, allowing clutter-induced spurious temporal responses to persist and undermine reliable detection. To address this issue, we propose motion-structure verification (MoSVer) for reliable learning, which introduces explicit joint verification of motion-derived temporal cues and structural cues to yield verified evidence for reliability-oriented supervision. Specifically, the motion-derived cue extraction (MDCE) module generates a motion support map to capture target-relevant motion evidence against the background. Meanwhile, the structure-constrained cue extraction (SCCE) module extracts a structural support map from regions exhibiting cross-frame structural consistency. Finally, the verification-guided matching (VGM) module integrates the two support maps to derive verified evidence, which is used as a reliability-aware matching prior to favor assignments supported by both cues. With RT-DETR as the base detector, MoSVer improves mAP ${}_{50:95}$ by 0.042 and 0.011 on airborne object tracking (AOT) and UAVSwarm, respectively, while reducing FP/frame@R = 0.50 by 17.8 % and 16.0 %. It also lowers MR@FP = 1 by 1.9 and 0.5 percentage points and decreases expected calibration error (ECE) by 0.025 and 0.009, providing quantitative evidence of improved detection reliability in clutter-dominated, low-SCR scenes.
Airborne infrared small-object tracking is crucial for applications such as autonomous reconnaissance and border surveillance. Unlike visible-light imagery, infrared data provides stable imaging in low-light and hazy conditions. However, tracking in this domain is exceptionally challenging due to the diminutive size of...
Ming-Yu Hong, Xue Jin, Yuan Liu et al.· IEEE Access· 1 citation
Small-object detection in unmanned aerial vehicle (UAV) remote-sensing imagery remains difficult because targets often have low spatial resolution, dense spatial distribution, partial occlusion and strong background clutter. These factors weaken discriminative features and restrict real-time inference on embedded edge...
Shuai-Jie Nie, Jia-Jian Yang, Xin He et al.· Engineering Research Express· 0 citations
A context-gated dynamic perception framework that treats small-object feature degradation as a coupled problem of representation, fusion, and prediction and indicates a practical accuracy-efficiency trade-off for dense aerial small-object perception.
Guang-Jun Gao, Ruibing Xie· Pattern Analysis and Applica...· 0 citations
Airborne optical tracking of uncrewed aerial vehicle (UAV) swarms is challenging due to extremely small target scales, rapid viewpoint changes, and cluttered backgrounds, which can weaken target feature responses and lead to intermittent or temporarily missing detector responses. Existing multi-object tracking methods...
Zhao-Chen Chu, Tao Song, Ren Jin et al.· 0 citations
Highlights What are the main findings? MDF-Det establishes a unified coarse-to-fine framework that jointly addresses weak-target preservation, dense-target decoupling, and scene-induced false-alarm suppression in wide-area motion imagery. The spatial attention-guided decoupling and learned scene-prior filtering mechani...
Kang Li, Xiao-Ran Zhang, Zheng Zhang et al.· Italian National Conference...· 0 citations
Joint detection-and-embedding (JDE) trackers avoid per-detection crop inference by reading identity features for re-identification (ReID) from the detector. The detection-center readout, however, does not use a track prediction when forming the appearance descriptor. This is restrictive in unmanned aerial vehicle (UAV)...