Collaborative Context-Affine Perception Network for Remote Sensing Small Object Detection
Abstract
Small object detection in remote sensing images (RSIs) is challenging because imaging degradation weakens object textures, reduces contrast, and blurs boundaries. These effects are further aggravated by hierarchical feature extraction, where repeated downsampling weakens shallow spatial cues before they reach deeper semantic representations. To alleviate these effects, we propose the collaborative context-affine perception network (CoCAPNet), which progressively refines structural, contextual, and spatial representations along the detection pipeline. CoCAPNet comprises four complementary components. First, the context-aware dynamic affine module (CDAM) separates feature responses into structure- and detail-dominant components and adaptively recombines them to retain weak spatial cues. Next, the dual-range perception fusion module (DPFM) adaptively balances local detail and global semantics by integrating scale-dependent response predictions with multilevel receptive field processing. The collaborative multikernel block (CMKB) then addresses sparse and unstable activations in small objects by enhancing object-relevant patterns via dynamic multikernel perception, fine-grained refinement, and channel-aware weighting. Finally, the fine-grained small object detector (FSDetector) introduces a high-resolution prediction branch to retain spatial detail for small targets. These components form a progressive feature-processing pipeline from representation reconstruction to scale-aware fusion, structural refinement, and high-resolution prediction. Experiments on the DIOR, NWPU VHR-10, and AI-TOD yield mean average precision (mAP) values of 86.3%, 95.8%, and 59.1%, respectively, indicating consistent gains across datasets with substantial scale variation and densely distributed small objects.