Pi-Noise-Guided Student–Teacher Distillation Detector for Open-Vocabulary Aerial Object Detection
Abstract
Open-vocabulary aerial object detection aims to detect objects outside the training sets, which typically involves distilling knowledge from pretrained vision-language models, e.g., RemoteCLIP, to inherit its generalizable recognition ability and thereby enabling the models to detect novel categories. However, existing distillation methods lack a customized perception incentive mechanism, leading to a disconnect between perception with discrimination knowledge and poor inductive generalization. To this end, we propose a positive-incentive-noise-guided progressive prior accumulation student–teacher distillation network, which utilizes the structured perturbation of noise distribution to achieve tighter localization–classification synergy. Concretely, we design an adaptive guided noise generator to continuously accumulate universal patterns of object-location distributions and refined semantic knowledge. These prior patterns not only provide implicit cues to simplify the detection task but also establish bidirectional knowledge flow, forming a closed-loop optimization of localization and classification. Then, we introduce a semantic-aware dual-alignment reprojection module, which employs cross-modal attention to achieve dual visual-semantic alignment and enhance feature discriminability. This module effectively mines the latent localization-aware information embedded in external teachers, thereby improving the model’s recognition sensitivity to unseen categories. Furthermore, to enhance data diversity and increase the number of pseudosamples, we propose a pixel-level adaptive CutMix strategy. This approach enriches training scenarios by performing pixel thresholding after channel separation. Extensive experiments on multiple remote sensing object detection benchmarks demonstrate that the proposed method achieves highly competitive performance compared with recent state-of-the-art methods, while maintaining a simple training pipeline without relying on additional classification or caption datasets.