Skip to content
Open access

MA-YOLO for robust multimodal feature fusion under weak cross-modal alignment in remote sensing object detection

Sep 2026 · Scientific Reports · 0 citations

Abstract

Multimodal remote sensing object detection benefits from the complementary characteristics of visible and infrared modalities. However, owing to the heterogeneous imaging mechanisms of RGB and infrared sensors, feature representations from the two modalities progressively become semantically inconsistent during hierarchical feature extraction, even when the input image pairs are geometrically registered. Such feature-level semantic inconsistency weakens the robustness of multimodal feature fusion and ultimately degrades object localization accuracy. To address this issue, we propose Misalignment-Aware YOLO (MA-YOLO). MA-YOLO incorporates a Global Semantic Encoder (GSE) to capture long-range contextual dependencies, a Robust Adaptive Fusion (RAF) module to adaptively refine multiscale feature representations by suppressing spatially inconsistent responses during hierarchical feature propagation, and an NWD-assisted bounding-box regression strategy to improve localization robustness under spatial uncertainty and ambiguous object boundaries. Extensive experiments on the VEDAI benchmark demonstrate that MA-YOLO achieves an average mAP $$_{50}$$ of 76.93% under tenfold cross-validation. Spatial perturbation experiments further show that the proposed framework maintains relatively stable detection performance under cross-modal shifts of up to 32 pixels, supporting its robustness to cross-modal spatial inconsistency. Additional experiments on the DIOR and NWPU VHR-10 datasets also demonstrate the applicability of the proposed framework to single-modal optical remote sensing scenarios.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.