A model-agnostic saliency-depth foreground-conditioning strategy combining appearance-based saliency with monocular relative depth to construct a coarse tower prior and suppress irrelevant content is proposed, improving robustness in cluttered scenes.
Abstract
Fine-grained segmentation of communication-tower components in UAV imagery is essential for automated inspection, yet task-specific models are hard to develop due to limited instance-level annotations. Zero-shot segmentation models offer a promising alternative, but in cluttered scenes, visually similar background structures interfere with component localization, causing missed instances and false positives. We propose a model-agnostic saliency-depth foreground-conditioning strategy combining appearance-based saliency with monocular relative depth to construct a coarse tower prior and suppress irrelevant content. We integrate this module with Grounded-SAM and SAM 3, yielding SD-Grounded-SAM and SD-SAM 3. SD-Grounded-SAM further applies geometric and depth-aware box refinement before mask generation, while SD-SAM 3 relies on SAM 3's internal setup. On TOW-300, a dataset of 340 communication-tower UAV images, our strategy improves both baselines: SD-SAM 3 achieves the strongest instance-segmentation performance, while SD-Grounded-SAM produces fewer false positives. Ablations confirm complementary gains from saliency, depth, and box refinement, improving robustness in cluttered scenes.
A context-gated dynamic perception framework that treats small-object feature degradation as a coupled problem of representation, fusion, and prediction and indicates a practical accuracy-efficiency trade-off for dense aerial small-object perception.
Guang-Jun Gao, Ruibing Xie· Pattern Analysis and Applica...· 0 citations
DLPANet is proposed, a novel dual-level prototype alignment network centered on Prototype-Guided Spatial Attention, enabling simultaneous modeling of scene context and fine-grained details and demonstrates that the decoupled dual cross-attention mechanism provides superior prototype-query alignment compared to prior gl...
Mustafa Alawadi, M. Fateh· Jordanian Journal of Compute...· 0 citations
Contrastive Language-Image Pre-training (CLIP) has demonstrated impressive capabilities in zero-shot transfer but often struggles with dense prediction tasks due to low spatial resolution and the loss of structural information. To address these limitations, we propose MARS-CLIP (Multi-resolution and Attention Refined S...
Nagito Saito, Shintaro Ito, Koichi Ito et al.· International Conference on...· 0 citations
HSAR-DETR is proposed, a detection framework that jointly improves hierarchical feature representation, cross-scale refinement, and geometry-aware localization and experimental results on the VisDrone, RSOD, and TinyPerson datasets demonstrate improved detection performance.
Pedestrian detection in crowds is a challenging problem in computer vision. Existing occlusion-handling methods heavily rely on expensive visible-box annotations to locate visible body parts, posing severe limitations in label acquisition cost and open-world generalization. To break through this limitation, we propose...
Tian Qiu, Jifeng Shen, Xin Zuo· Italian National Conference...· 0 citations
Small-object detection in unmanned aerial vehicle (UAV) imagery remains challenging because target objects often occupy only a few pixels, exhibit weak feature responses, and are easily obscured by complex backgrounds. These aspects significantly limit the effectiveness of end-to-end detection systems. To overcome thes...