A CLIP-Guided Progressive Body-Part Semantic Alignment Network is proposed, termed PBSA-Net, which introduces CLIP-derived textual semantics as modality-agnostic guidance for both global representation learning and local body-part feature extraction.
Abstract
Visible-infrared person re-identification (VI-ReID) aims to retrieve pedestrian images of the same identity across visible and infrared modalities, but remains challenging due to the large modality gap and unstable local correspondence. Existing methods mainly rely on visual cues, which may be insufficient when infrared images lack color and fine-grained texture information. To address this issue, this paper proposes a CLIP-Guided Progressive Body-Part Semantic Alignment Network, termed PBSA-Net. The proposed method introduces CLIP-derived textual semantics as modality-agnostic guidance for both global representation learning and local body-part feature extraction. Specifically, a global semantic branch first learns identity-level textual anchors to regularize global visual features. Then, a body-part semantic branch exploits identity-aware body-part prompt learning, multi-level feature fusion, and text-guided cross-attention to guide fine-grained local representation learning. A progressive three-stage optimization strategy is further adopted to decouple global semantic learning, body-part semantic correspondence learning, and retrieval-oriented feature optimization. Experiments on SYSU-MM01, RegDB, and LLCM demonstrate the effectiveness of PBSA-Net. It achieves 76.5% Rank-1 and 74.2% mAP on SYSU-MM01, 82.5% Rank-1 and 76.0% mAP on RegDB, and 61.8% Rank-1 and 65.8% mAP on LLCM. Ablation studies further show that the proposed body-part semantic alignment and progressive optimization provide complementary improvements.
It is argued that VI-ReID should be treated as an early cross-modal correspondence learning problem rather than only a late embedding alignment problem, and CMIA-Net is proposed, a framework that establishes bidirectional visible-infrared interaction at shallow backbone stages and introduces Spectral-Invariant Augmenta...
Dao-Li Zhang, Qi-Cheng Liu· Engineering Research Express· 0 citations
Structural-Semantic Reciprocal Learning (SSRL), a framework that transforms open-loop association into a self-correcting closed-loop system, achieves robust cross-modal representation through the reciprocal interaction between structural and semantic learning.
Moyao Tian, Shijia Liu, Yan Yang et al.· arXiv.org· 0 citations
A Multi-level Semantic-Guided (MSG) framework that integrates contextual and fine-grained visual information to eliminate clothing variance across both conceptual and pixel dimensions is proposed.
Shijuan Huang, Hefei Ling, Zongyi Li et al.· ACM Transactions on Multimed...· 0 citations
A Progressively Biased Split Vision Transformer (PBSVT) is proposed, which combines a split ViT backbone with progressive bias training to gradually reduce RGB-dominant bias while preserving modality-shared structure and demonstrates the effectiveness of progressive modality transition for robust VI-ReID representation...
Mengru Jiao, Xin-Yue Xu, Jun-Feng Zhang· International journal of pat...· 0 citations
This work proposes CLIP-SGI, a semantic-guided and instance-consistent framework for generalizable person ReID that combines semantic guidance, domain-aware representation learning, and instance consistency to improve robustness under domain shifts.
Dai-Xin Liu, Yu Yang, Linlin Tang et al.· IEEE Transactions on Image P...· 0 citations
Visible-infrared person re-identification (VI-ReID) is an important technique for around-the-clock person matching, and its primary challenge arises from substantial cross-modal discrepancies. To address this challenge, we propose a Decoupled Information-Guided Cross-Modal Alignment (DIGCA) framework organized into thr...