2026· IEEE Transactions on Geoscience and Remote Sensing· Vol 64, pp. 5408215-5408215· 0 citations· 30 references
Abstract
Multimodal object detection that integrates visible-spectrum red-green-blue (RGB) and infrared (IR) imagery has become increasingly important for robust perception in challenging conditions such as low illumination, cluttered backgrounds, and small targets, especially in uncrewed aerial vehicle (UAV) and edge-computing scenarios. However, many existing RGB–IR detection frameworks rely on heavy backbones and fixed or aggressive fusion strategies, which introduce redundant computation and uncontrolled cross-modal noise, limiting their applicability to real-time and resource-constrained platforms. To address these limitations, this article proposes an efficient and lightweight RGB–IR dual-stream detection framework that emphasizes adaptive cross-modal interaction. The proposed method adopts a dual-stream backbone to preserve modality-specific representations and introduces adaptive gated cross-modal interaction (AGCMI) at early stages to enable learnable gated feature interaction, regulating information exchange between RGB and IR branches for effective noise suppression. At later stages, a weight-guided aggregation (WGA) module performs two-stage adaptive fusion with a residual-dominant design, promoting stable performance with minimal parameter overhead while reducing the risk of performance degradation. In addition, a lightweight partial-channel C2f (PC2f) backbone variant is employed to further reduce parameters and computational overhead. Experiments conducted on the DroneVehicle dataset demonstrate that the proposed model achieves 81.6% mAP on the test set with only 4.1M parameters. Compared with representative state-of-the-art (SOTA) multimodal detectors, the proposed approach attains competitive accuracy while offering a substantial reduction in model size and improved efficiency. These results demonstrate that the proposed RGB–IR dual-stream detection framework achieves an effective balance between detection performance and deployment efficiency, suggesting strong potential for real-time multimodal perception tasks on resource-constrained platforms.
RGB-infrared (RGB-IR) fusion improves object detection by enabling robust object localization under challenging illumination and environmental conditions. However, the additional IR modality often increases computational cost, limiting its deployment in real-time measurement and perception systems. This work revisits c...
Hongru Xiao, Bo Li, Bin Yang et al.· Measurement science and tech...· 0 citations
Small-object detection under low illumination remains a persistent challenge in aerial surveillance, nighttime patrol, and safety-critical vision tasks, where single-modality sensors—RGB or infrared alone—provide degraded or incomplete target information. This paper proposes an RGB-T object detection network that combi...
Experimental results show that ZOS-Net achieves a favorable balance between detection performance and complexity on the M3FD, FLIR-aligned, and VEDAI datasets and provides a lightweight solution for RGB-T object detection on resource-constrained platforms.
Object detection using visible–infrared images has become increasingly important for all-day detection scenarios. However, due to significant imaging discrepancies between the visible and infrared modalities, achieving accurate modal alignment and effective feature fusion remains a major challenge. Existing methods oft...
Kai-Yue Men, Cheng-You Wang, Xiao Zhou et al.· IEEE Transactions on Geoscie...· 0 citations
Object detection in complex lighting and harsh environments significantly benefits from the synergistic deployment of visible and infrared spectra. However, most existing dual-modality methods follow a two-stage pipeline termed "fusion-then-detection", inevitably entailing substantial model complexity and intensive c...
Chao Zeng, Hao Zhao, Ju Zhou· Measurement science and tech...· 0 citations
A Boundary-Guided Dual-Perspective Cross-Modal Fusion Network (BDPNet) is proposed to explicitly preserve shallow geometric structures and decouple deep semantic fusion into macroscopic and microscopic perspectives.
Hu Lin, Zhi-Wei Fu, Xiu-Mei Chen et al.· Remote Sensing· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.