CAF-YOLO: YOLOv8-based Bimodal Pedestrian Detection Method Using Cross-Attention Fusion
Abstract
Visible-infrared bimodal pedestrian detection holds significant application value in scenarios such as intelligent surveillance and autonomous driving. However, visible images suffer performance degradation under low-light conditions, while infrared images lack detailed texture information. Moreover, existing fusion methods still have limitations in real-time performance, interaction efficiency, and lightweight design. To address these challenges, this paper proposes an improved YOLOv8 model based on cross-attention fusion, named CAF-YOLO. The method designs a Cross-Attention Fusion (CAF) module to achieve feature complementarity and semantic alignment between the two modalities through bidirectional attention mechanism. Meanwhile, an Efficient Multi-scale Attention (EMA) module is introduced to enhance feature representation capability. Experiments on LLVIP and FLIR datasets show that CAF-YOLO achieves mAP@0.5 improvements of 9.1% and 1.2% over visible and infrared single-modal baselines on LLVIP, and 17.0% and 3.6% on FLIR, respectively. The model has only 11.74M parameters, significantly outperforming fusion methods with higher computational complexity. Ablation studies and visualization analysis verify the effectiveness of each module, and the model maintains robust detection capability in low-light and occlusion scenarios.