This work proposes MDCRNet, a Multi-scale Decomposed Convolution Refinement Network that enhances cross-modal feature learning and discriminative metric learning, and develops a Joint Discriminative Metric Loss incorporating a novel Granularity Discriminative Loss (GDL).
Abstract
Visible-infrared person re-identification (VI-ReID) suffers from cross-modal discrepancies and limited discriminative capabilities, leading to suboptimal recognition performance. Current approaches exhibit limitations in semantic mining, cross-modal fusion and feature constraints. To tackle these challenges, we propose MDCRNet, a Multi-scale Decomposed Convolution Refinement Network that enhances cross-modal feature learning and discriminative metric learning. Specifically, we introduce a Hierarchical Learning Module (HLM) containing four Hierarchical Decomposed Convolution Attention (HDCA) modules, each equipped with lightweight channel attention and multi-scale spatial perception blocks to capture multi-scale spatial dependencies. Moreover, we develop a Joint Discriminative Metric Loss (JDML) incorporating a novel Granularity Discriminative Loss (GDL) that simultaneously optimizes intra-identity compactness and inter-identity separability across modalities. Extensive experiments on SYSU-MM01 and RegDB datasets demonstrate that MDCRNet achieves state-of-the-art performance on both benchmarks. Code is available at https://github.com/Kevin-zms/MDCRNet.
This study proposes a novel dual-path convolution based multi-scale feature alignment (DCMFA) network that significantly outperforms existing mainstream methods in terms of recognition accuracy.
Bailiang Huang, Bin Chen, Tian-Ran Sun et al.· 電腦學刊· 0 citations
Cross-modal person re-identification between visible and infrared domains remains a challenging problem due to significant modality gaps. This paper presents a novel approach termed Intermediate Shared Feature Network (ISFNet) that explicitly addresses this issue by exploiting intermediate feature representations withi...
Aobo Fan, Wangmeng Wang, Zhixin Tie et al.· Electronics· 0 citations
A Progressively Biased Split Vision Transformer (PBSVT) is proposed, which combines a split ViT backbone with progressive bias training to gradually reduce RGB-dominant bias while preserving modality-shared structure and demonstrates the effectiveness of progressive modality transition for robust VI-ReID representation...
Mengru Jiao, Xin-Yue Xu, Jun-Feng Zhang· International journal of pat...· 0 citations
Visible-infrared person re-identification (VI-ReID) is an important technique for around-the-clock person matching, and its primary challenge arises from substantial cross-modal discrepancies. To address this challenge, we propose a Decoupled Information-Guided Cross-Modal Alignment (DIGCA) framework organized into thr...
It is argued that VI-ReID should be treated as an early cross-modal correspondence learning problem rather than only a late embedding alignment problem, and CMIA-Net is proposed, a framework that establishes bidirectional visible-infrared interaction at shallow backbone stages and introduces Spectral-Invariant Augmenta...
Dao-Li Zhang, Qi-Cheng Liu· Engineering Research Express· 0 citations
Person re-identification (Re-ID) aims to retrieve the same person across non-overlapping cameras. Despite recent progress, Re-ID remains highly challenging due to high inter-class similarity in large-scale datasets and drastic intra-class variations in cross-platform scenarios (e.g., drones and wearable cameras). While...
Meifeng Liu, Hua Han, A. A. M. Muzahid et al.· IEEE Access· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.