Skip to content
Open access

Boundary-Guided Dual-Perspective Cross-Modal Fusion Network for RGB-IR Object Detection

Sep 2026 · Remote Sensing · Vol 18, pp. 3175 · 0 citations · 35 references

TL;DR

A Boundary-Guided Dual-Perspective Cross-Modal Fusion Network (BDPNet) is proposed to explicitly preserve shallow geometric structures and decouple deep semantic fusion into macroscopic and microscopic perspectives.

Abstract

Visible-infrared (RGB-IR) object detection leverages multimodal information to ensure reliable perception in complex environments. However, dynamic scenes pose significant challenges due to the frequent inconsistency between scene-level modality contributions and local spatial reliability. Furthermore, standard feature extraction progressively attenuates boundary-sensitive structural cues, and unified fusion strategies often fail to capture spatially varying cross-modal complementarity. To overcome these limitations, we propose a Boundary-Guided Dual-Perspective Cross-Modal Fusion Network (BDPNet) to explicitly preserve shallow geometric structures and decouple deep semantic fusion into macroscopic and microscopic perspectives. Specifically, a Geometric Boundary Enhancement Module (GBEM) embeds Sobel-based high-frequency priors into shallow dual-modal features via residual spatial modulation, preventing the loss of crucial localization cues during downsampling. In the deep semantic space, a Hybrid Dual-Perspective Adaptive Fusion Module (HDAM) employs an illumination-aware branch for global modality weighting and a spatial confidence-driven branch for local cross-modal rectification. A spatial gating mechanism then dynamically reconciles these macro-environmental and micro-signal features. Extensive experiments on M3FD, LLVIP, and DroneVehicle demonstrate the effectiveness of BDPNet. Compared with state-of-the-art methods, BDPNet improves mAP50–95 by 0.8% and 1.0% on M3FD and LLVIP, respectively, and improves mAP50 by 0.6% on DroneVehicle, while using substantially fewer parameters and lower computational cost.

Read PDF

Similar papers

Preprint Aug 2026

Bridging Severe Cross-Modal Misalignment: End-to-End Visible-Infrared Object Detection via Explicit Feature-Domain Affine Registration

Visible-infrared object detection relies on complementary RGB and thermal cues, but its performance is often degraded by cross-modal spatial misalignment. Most existing methods rely on implicit feature adaptation to handle weakly misaligned scenarios, while large-offset geometric discrepancies remain insufficiently add...

Ming Qi, Yuyang Wang, Ming-Jing Zhao et al. · 0 citations
2026

MAMENet: Modal Alignment and Multiscale Feature Enhancement Network for RGB–IR Small Object Detection

Object detection using visible–infrared images has become increasingly important for all-day detection scenarios. However, due to significant imaging discrepancies between the visible and infrared modalities, achieving accurate modal alignment and effective feature fusion remains a major challenge. Existing methods oft...

Kai-Yue Men, Cheng-You Wang, Xiao Zhou et al. · 0 citations
2026

D2CMFDet: Disparity-Guided Dynamic Cross-Modal Mamba Fusion for Multispectral Object Detection

Existing multispectral object detectors struggle to balance wide-area cross-modal discrepancy modeling with high-resolution local detail preservation. To address this bottleneck, we propose D2CMFDet, a disparity-guided dynamic fusion framework that integrates learnable state-transition modeling with cross-modal feature...

Hai Yang, Xing-Du Wu, Chen-Hai Wei et al. · 0 citations
Preprint Aug 2026

EGM-Det: Entropy-Guided Multimodal Adaptive Fusion for UAV RGB-IR Object Detection

Joint use of RGB and infrared (IR) imagery can improve UAV-view object detection, but most existing methods fuse multimodal features with static or fixed weights and therefore overlook spatially varying modality reliability. We propose EGM-Det, an entropy-guided multimodal adaptive fusion framework for RGB-IR object de...

Cun-Zheng Fan, Dawei Yan, Guan-Lin Wang et al. · 0 citations
Open access Sep 2026

MA-YOLO for robust multimodal feature fusion under weak cross-modal alignment in remote sensing object detection

Multimodal remote sensing object detection benefits from the complementary characteristics of visible and infrared modalities. However, owing to the heterogeneous imaging mechanisms of RGB and infrared sensors, feature representations from the two modalities progressively become semantically inconsistent during hie...

Jian-Qiong Huang, Zhi-Hong Lin, Waqar Khan et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.