An Adaptive Network for Enhanced Vision and Semantic Cooperative Understanding in Aerial Object Detection
Abstract
Object detection in remote sensing imagery faces challenges such as extreme scale variations and complex backgrounds. Although current methods have made significant strides in visual feature extraction, their predominant focus remains on the image itself, overlooking the potential of integrating external knowledge. To address this limitation, we introduce the knowledge-aware network with region-adaptive fusion for detection (KARFDet), which seamlessly integrates region-specific semantic information with visual features. First, a multiscale fused kernel attention (MSFKA) module is introduced, leveraging a parallel multibranch architecture to enhance contextual feature extraction. Second, a knowledge graph semantic extraction (KGSE) module is designed, employing the random walk with restart (RWR) algorithm to transform discrete knowledge into computable semantic associations. Finally, a novel triple-order knowledge integration (TOKI) mechanism is proposed, which adaptively fuses original, second-order, and probabilistic semantic knowledge, dynamically allocating knowledge weights based on target scale characteristics. Experiments on the DIOR, NWPU VHR-10, and SIMD datasets show that KARFDet achieves mAP50 scores of 65.6%, 91.9%, and 75.8%, respectively, significantly outperforming the baseline model and establishing a new paradigm for semantic-aware detection in complex scenarios. The code is available at https://github.com/ChengXCode/KARFDet