Skip to content

MAMENet: Modal Alignment and Multiscale Feature Enhancement Network for RGB–IR Small Object Detection

2026 · IEEE Transactions on Geoscience and Remote Sensing · Vol 64, pp. 5408815-5408815 · 0 citations · 50 references

Abstract

Object detection using visible–infrared images has become increasingly important for all-day detection scenarios. However, due to significant imaging discrepancies between the visible and infrared modalities, achieving accurate modal alignment and effective feature fusion remains a major challenge. Existing methods often perform inadequately in handling modal mismatch and feature fusion representation, resulting in unstable detection performance. To address these challenges, we propose a novel modal alignment and multiscale feature enhancement network (MAMENet). In particular, we first design a cross-modal alignment and interaction (CMAI) module, which combines deformable convolution with a cross-attention mechanism to guide feature alignment and deep interaction between visible and infrared modalities. Second, we introduce a multiscale feature enhancement (MSFE) module to capture rich contextual information through a multiscale dilation strategy, thereby enhancing feature representations. Furthermore, an adaptive fusion (AF) module is proposed to dynamically assign fusion weights, achieving more reliable and flexible cross-modal fusion. To demonstrate the superiority of the proposed method, extensive comparative experiments and ablation studies are conducted on three public visible–infrared datasets, namely VEDAI, FLIR, and M3FD. The proposed MAMENet achieves mAP50 scores of 88.5%, 87.2%, and 87.7% on the three datasets, respectively, outperforming other state-of-the-art methods.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.