Skip to content
Preprint

Saliency-Depth Conditioning for Zero-Shot Segmentation of Communication-Tower Components in Cluttered UAV Imagery

Aug 2026 · 0 citations · 43 references
Computer Science

TL;DR

A model-agnostic saliency-depth foreground-conditioning strategy combining appearance-based saliency with monocular relative depth to construct a coarse tower prior and suppress irrelevant content is proposed, improving robustness in cluttered scenes.

Abstract

Fine-grained segmentation of communication-tower components in UAV imagery is essential for automated inspection, yet task-specific models are hard to develop due to limited instance-level annotations. Zero-shot segmentation models offer a promising alternative, but in cluttered scenes, visually similar background structures interfere with component localization, causing missed instances and false positives. We propose a model-agnostic saliency-depth foreground-conditioning strategy combining appearance-based saliency with monocular relative depth to construct a coarse tower prior and suppress irrelevant content. We integrate this module with Grounded-SAM and SAM 3, yielding SD-Grounded-SAM and SD-SAM 3. SD-Grounded-SAM further applies geometric and depth-aware box refinement before mask generation, while SD-SAM 3 relies on SAM 3's internal setup. On TOW-300, a dataset of 340 communication-tower UAV images, our strategy improves both baselines: SD-SAM 3 achieves the strongest instance-segmentation performance, while SD-Grounded-SAM produces fewer false positives. Ablations confirm complementary gains from saliency, depth, and box refinement, improving robustness in cluttered scenes.

View source

Similar papers

Aug 2026

Context-gated dynamic perception for small-object detection in dense aerial scenes

A context-gated dynamic perception framework that treats small-object feature degradation as a coupled problem of representation, fusion, and prediction and indicates a practical accuracy-efficiency trade-off for dense aerial small-object perception.

Guang-Jun Gao, Ruibing Xie · 0 citations
Open access 2026

Dual-Level Prototype Alignment via Cross-Attention for Few-Shot Remote Sensing Semantic Segmentation

DLPANet is proposed, a novel dual-level prototype alignment network centered on Prototype-Guided Spatial Attention, enabling simultaneous modeling of scene context and fine-grained details and demonstrates that the decoupled dual cross-attention mechanism provides superior prototype-query alignment compared to prior gl...

Mustafa Alawadi, M. Fateh · 0 citations
Conference Open access Sep 2026

MARS-CLIP: Multi-Resolution and Attention Refined Zero-Shot Image Segmentation

Contrastive Language-Image Pre-training (CLIP) has demonstrated impressive capabilities in zero-shot transfer but often struggles with dense prediction tasks due to low spatial resolution and the loss of structural information. To address these limitations, we propose MARS-CLIP (Multi-resolution and Attention Refined S...

Nagito Saito, Shintaro Ito, Koichi Ito et al. · 0 citations
Open access Aug 2026

HSAR-DETR: Hierarchical Spatial–Frequency Attention Network for UAV Small Object Detection

HSAR-DETR is proposed, a detection framework that jointly improves hierarchical feature representation, cross-scale refinement, and geometry-aware localization and experimental results on the VisDrone, RSOD, and TinyPerson datasets demonstrate improved detection performance.

Cheng Zhang, Zhibo Guo · 0 citations
Open access Aug 2026

STG: Structured Topology of Gridpoints for Occluded Pedestrian Detection

Pedestrian detection in crowds is a challenging problem in computer vision. Existing occlusion-handling methods heavily rely on expensive visible-box annotations to locate visible body parts, posing severe limitations in label acquisition cost and open-world generalization. To break through this limitation, we propose...

Tian Qiu, Jifeng Shen, Xin Zuo · 0 citations
Open access Aug 2026

MPC-DETR: A Multi-Scale Patch Context Transformer for Small Object Detection in UAV Imagery

Small-object detection in unmanned aerial vehicle (UAV) imagery remains challenging because target objects often occupy only a few pixels, exhibit weak feature responses, and are easily obscured by complex backgrounds. These aspects significantly limit the effectiveness of end-to-end detection systems. To overcome thes...

Quan-Xiang Wang, Zhao-Fa Zhou, Zhi-Li Zhang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.