This study proposes a novel dual-path convolution based multi-scale feature alignment (DCMFA) network that significantly outperforms existing mainstream methods in terms of recognition accuracy.
Abstract
To address the performance degradation of person re-identification (ReID) under complex lighting and day-night conditions, this study proposes a novel dual-path convolution based multi-scale feature alignment (DCMFA) network. The network mainly focuses on addressing the challenges of modality discrepancy and feature alignment between visible and infrared images. First, to cope with feature deformation in persons caused by external factors in both modalities, we design a dual-path convolution (DPCon) layer. Secondly, to further alleviate feature loss caused by scale variation, we construct a multi-scale feature aggregation (MSFA) module by stacking DPCon layers of different depths and incorporating an attention mechanism to effectively aggregate key multi-scale information and suppress redundancy. Finally, to enable more effective alignment of multi-scale features across the two modalities, we propose a feature mapping and alignment operations (FMAO) module. Experimental results on three publicly available cross-modal visible-infrared person re-identification (VI-ReID) datasets demonstrate that our DCMFA network significantly outperforms existing mainstream methods in terms of recognition accuracy. Specifically, on the SYSU-MM01 dataset, our method achieves a Rank-1 accuracy of 83.74% in the single-shot setting and 89.47% in the multi-shot setting.
This work proposes MDCRNet, a Multi-scale Decomposed Convolution Refinement Network that enhances cross-modal feature learning and discriminative metric learning, and develops a Joint Discriminative Metric Loss incorporating a novel Granularity Discriminative Loss (GDL).
Mingsheng Zheng, Zirui Jiang, Bo Liu et al.· 0 citations
Visible-infrared person re-identification (VI-ReID) is an important technique for around-the-clock person matching, and its primary challenge arises from substantial cross-modal discrepancies. To address this challenge, we propose a Decoupled Information-Guided Cross-Modal Alignment (DIGCA) framework organized into thr...
It is argued that VI-ReID should be treated as an early cross-modal correspondence learning problem rather than only a late embedding alignment problem, and CMIA-Net is proposed, a framework that establishes bidirectional visible-infrared interaction at shallow backbone stages and introduces Spectral-Invariant Augmenta...
Dao-Li Zhang, Qi-Cheng Liu· Engineering Research Express· 0 citations
Cross-modal person re-identification between visible and infrared domains remains a challenging problem due to significant modality gaps. This paper presents a novel approach termed Intermediate Shared Feature Network (ISFNet) that explicitly addresses this issue by exploiting intermediate feature representations withi...
Aobo Fan, Wangmeng Wang, Zhixin Tie et al.· Electronics· 0 citations
Object detection using visible–infrared images has become increasingly important for all-day detection scenarios. However, due to significant imaging discrepancies between the visible and infrared modalities, achieving accurate modal alignment and effective feature fusion remains a major challenge. Existing methods oft...
Kai-Yue Men, Cheng-You Wang, Xiao Zhou et al.· IEEE Transactions on Geoscie...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.