Text Semantic Prior Contrastive Learning for Few-Shot Remote Sensing Object Detection
Abstract
Deep-learning-based detectors for remote sensing object detection have achieved remarkable success, but their performance heavily depends on large-scale annotated datasets, which are costly and time-consuming. Few-shot remote sensing object detection has, therefore, emerged as a promising paradigm to recognize novel categories using only a few labeled samples, where transfer-learning-based methods can achieve excellent performance via simple fine-tuning. However, high interclass similarity still poses a significant challenge, leading to misidentification among confusing categories. To address this problem, a text semantic prior few-shot object detection network (TSP-Net) is proposed for few-shot remote sensing object detection. A remote sensing object characteristic prompt is designed, and a CLIP text encoder is used to obtain class-level semantic relationships, which serve as the text semantic prior. This prior is then coupled with contrastive learning in the detection network, where interclass feature repulsion is adaptively weighted according to semantic similarity, thereby enlarging distance among confusing categories. Integrating text semantic prior with contrastive learning in a unified detection network, TSP-Net effectively enhances feature discrimination, improving detection performance. Extensive experiments on the DIOR and NWPU VHR-10 benchmarks demonstrate that the proposed method consistently improves novel-class detection performance, achieving notable gains under challenging three-shot and five-shot settings.