Semantic-Aware Consistency Rectification for Cross-Modal Person Retrieval
Text-to-Image Person Retrieval (TIPR) faces significant challenges due to the “one-to-many” nature of cross-modal matching, where a single identity corresponds to multiple images with varying viewpoints. Standard metric learning approaches often enforce rigid alignment between a text description and all same-identity images, ignoring the semantic asymmetry between strictly paired and unpaired samples. This indiscriminate “hard-labeling” introduces optimization noise and fails to account for the lexical ambiguity inherent in natural language. To overcome these limitations, we propose the Dynamic Calibration and Robust Distillation Network (DCR-Net). Our approach introduces a Paired-Guided Consistency Rectification (PCR) module that utilizes strictly paired data as reliable anchors to dynamically calibrate the confidence of unpaired samples, effectively suppressing noise from inconsistent associations under mild-to-moderate corruptions. Furthermore, we design a Semantic Knowledge Distillation (SKD) strategy incorporating a Token-level Resilience Modeling (TRM) task. By leveraging soft probability distributions from a momentum teacher, this mechanism enhances the model’s robustness against linguistic variations and synonymous descriptions. Extensive experiments on benchmarks demonstrate that DCR-Net achieves state-of-the-art performance by exhibiting strong resilience against semantic noise and alignment rigidity in controlled settings.