Aug 2026· International journal of pattern recognition and artificial intelligence· 0 citations
TL;DR
A Progressively Biased Split Vision Transformer (PBSVT) is proposed, which combines a split ViT backbone with progressive bias training to gradually reduce RGB-dominant bias while preserving modality-shared structure and demonstrates the effectiveness of progressive modality transition for robust VI-ReID representation learning.
Abstract
Visible-Infrared Person Re-Identification (VI-ReID) remains a challenging task due to the significant modality gap between visible and infrared images, which hinders accurate cross-modality identity matching. Existing methods often struggle to balance modality invariance and feature discriminability. Most methods employ ImageNet-pretrained backbones that are heavily biased toward RGB statistics, causing the extracted cross-modal features to over-emphasize visible-spectrum information and weakening the understanding of infrared cues. To address this issue, some works introduce third-modality generation or grayscale-based augmentation, but these strategies either increase training complexity or still leave a non-negligible discrepancy from real infrared data. We propose a Progressively Biased Split Vision Transformer (PBSVT), which combines a split ViT backbone with progressive bias training to gradually reduce RGB-dominant bias while preserving modality-shared structure. Extensive experiments on SYSU-MM01 and RegDB show that PBSVT achieves state-of-the-art or highly competitive performance while introducing no additional inference cost. PBSVT obtains the best results on 8 of the 12 reported indicators, including 77.90% Rank-1 and 97.98% Rank-10 on SYSU-MM01 all-search, 83.74% Rank-1 on SYSU-MM01 indoor-search, and 92.89% Rank-1, 98.77% Rank-10, and 92.18% mAP on RegDB Visible-to-Infrared. These results demonstrate the effectiveness of progressive modality transition for robust VI-ReID representation learning.
This work proposes MDCRNet, a Multi-scale Decomposed Convolution Refinement Network that enhances cross-modal feature learning and discriminative metric learning, and develops a Joint Discriminative Metric Loss incorporating a novel Granularity Discriminative Loss (GDL).
Mingsheng Zheng, Zirui Jiang, Bo Liu et al.· 0 citations
It is argued that VI-ReID should be treated as an early cross-modal correspondence learning problem rather than only a late embedding alignment problem, and CMIA-Net is proposed, a framework that establishes bidirectional visible-infrared interaction at shallow backbone stages and introduces Spectral-Invariant Augmenta...
Dao-Li Zhang, Qi-Cheng Liu· Engineering Research Express· 0 citations
Infrared and visible image fusion is pivotal for robust visual perception across all weather conditions and scenes. Although deep learning-based methods have made notable progress, most either assume pre-aligned inputs or rely on implicit feature-space alignment, which fails to fundamentally address the amplification o...
Jin-Yuan Liu, Zengxi Zhang, Jiahao Zhang et al.· IEEE Transactions on Pattern...· 1 citation
A Transformer-based baseline framework for visible-infrared ReID is proposed, designed to effectively capture modality-invariant features and outline several promising directions for future research.
Xiao Wang, Bing Wang, Bin Yang et al.· arXiv.org· 0 citations
Unsupervised person re-identification (USL-ReID) typically relies on clustering to generate pseudo-labels, but significant cross-view appearance variations often cause images of the same identity to be split into different clusters. Training on such noisy pseudo-labels severely degrades the learned representations. The...
Xuan Tan, Qi-Xian Zhang, Ding Qi et al.· IEEE Transactions on Image P...· 0 citations
Person re-identification (Re-ID) aims to retrieve the same person across non-overlapping cameras. Despite recent progress, Re-ID remains highly challenging due to high inter-class similarity in large-scale datasets and drastic intra-class variations in cross-platform scenarios (e.g., drones and wearable cameras). While...
Meifeng Liu, Hua Han, A. A. M. Muzahid et al.· IEEE Access· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.