A Boundary-Guided Dual-Perspective Cross-Modal Fusion Network (BDPNet) is proposed to explicitly preserve shallow geometric structures and decouple deep semantic fusion into macroscopic and microscopic perspectives.
Abstract
Visible-infrared (RGB-IR) object detection leverages multimodal information to ensure reliable perception in complex environments. However, dynamic scenes pose significant challenges due to the frequent inconsistency between scene-level modality contributions and local spatial reliability. Furthermore, standard feature extraction progressively attenuates boundary-sensitive structural cues, and unified fusion strategies often fail to capture spatially varying cross-modal complementarity. To overcome these limitations, we propose a Boundary-Guided Dual-Perspective Cross-Modal Fusion Network (BDPNet) to explicitly preserve shallow geometric structures and decouple deep semantic fusion into macroscopic and microscopic perspectives. Specifically, a Geometric Boundary Enhancement Module (GBEM) embeds Sobel-based high-frequency priors into shallow dual-modal features via residual spatial modulation, preventing the loss of crucial localization cues during downsampling. In the deep semantic space, a Hybrid Dual-Perspective Adaptive Fusion Module (HDAM) employs an illumination-aware branch for global modality weighting and a spatial confidence-driven branch for local cross-modal rectification. A spatial gating mechanism then dynamically reconciles these macro-environmental and micro-signal features. Extensive experiments on M3FD, LLVIP, and DroneVehicle demonstrate the effectiveness of BDPNet. Compared with state-of-the-art methods, BDPNet improves mAP50–95 by 0.8% and 1.0% on M3FD and LLVIP, respectively, and improves mAP50 by 0.6% on DroneVehicle, while using substantially fewer parameters and lower computational cost.
Visible-infrared object detection relies on complementary RGB and thermal cues, but its performance is often degraded by cross-modal spatial misalignment. Most existing methods rely on implicit feature adaptation to handle weakly misaligned scenarios, while large-offset geometric discrepancies remain insufficiently add...
Ming Qi, Yuyang Wang, Ming-Jing Zhao et al.· 0 citations
Object detection using visible–infrared images has become increasingly important for all-day detection scenarios. However, due to significant imaging discrepancies between the visible and infrared modalities, achieving accurate modal alignment and effective feature fusion remains a major challenge. Existing methods oft...
Kai-Yue Men, Cheng-You Wang, Xiao Zhou et al.· IEEE Transactions on Geoscie...· 0 citations
Existing multispectral object detectors struggle to balance wide-area cross-modal discrepancy modeling with high-resolution local detail preservation. To address this bottleneck, we propose D2CMFDet, a disparity-guided dynamic fusion framework that integrates learnable state-transition modeling with cross-modal feature...
Hai Yang, Xing-Du Wu, Chen-Hai Wei et al.· IEEE Transactions on Geoscie...· 0 citations
Joint use of RGB and infrared (IR) imagery can improve UAV-view object detection, but most existing methods fuse multimodal features with static or fixed weights and therefore overlook spatially varying modality reliability. We propose EGM-Det, an entropy-guided multimodal adaptive fusion framework for RGB-IR object de...
Cun-Zheng Fan, Dawei Yan, Guan-Lin Wang et al.· 0 citations
Multimodal remote sensing object detection benefits from the complementary characteristics of visible and infrared modalities. However, owing to the heterogeneous imaging mechanisms of RGB and infrared sensors, feature representations from the two modalities progressively become semantically inconsistent during hie...