Skip to content

D2CMFDet: Disparity-Guided Dynamic Cross-Modal Mamba Fusion for Multispectral Object Detection

2026 · IEEE Transactions on Geoscience and Remote Sensing · Vol 64, pp. 5408618-5408618 · 0 citations · 60 references

Abstract

Existing multispectral object detectors struggle to balance wide-area cross-modal discrepancy modeling with high-resolution local detail preservation. To address this bottleneck, we propose D2CMFDet, a disparity-guided dynamic fusion framework that integrates learnable state-transition modeling with cross-modal feature calibration. Here, disparity denotes a channelwise feature-response discrepancy between RGB and infrared branches in the learned representation space, rather than geometric or stereo disparity, and it is not an estimate of modality reliability. At its core lies a unified Disparity-Guided Cross-Modal Mamba Fusion (DGCM) engine, which follows a cascaded enhance-then-fuse architecture. Within DGCM, the dynamic feature enhancement module (DFEM) embeds a diagonal local differential convolution (DLDConv), implemented as a structurally modulated convolution with a diagonal bias, to strengthen direction-sensitive local structural cues before global fusion. After this local enhancement stage, the calibrated modality features are converted into a constructed joint feature, which is processed by the Joint-Feature Mamba Fusion (JFMF) module through linear-complexity state-space modeling. This design explicitly separates two roles: DFEM preserves modality-specific local calibration, whereas JFMF performs efficient global modeling on the constructed joint feature. Extensive experiments on public RGBT datasets demonstrate that D2CMFDet achieves state-of-the-art or highly competitive performance in detection accuracy, inference speed, and robustness with favorable computational overhead. Code is available at https://github.com/jacksonwu09/D2CMFDet

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.