Skip to content
Conference Open access

Re-Weighting Cross-Modal Pairs via Rank Consistency for Noise-Robust Retrieval

Sep 2026 · Proceedings of the Thirty-Fifth International Joint Conference on Artificial Intelligence · pp. 4705-4713 · 0 citations · 38 references

TL;DR

A novel semantic matching degree estimation method based on ranking consistency that reweights the training paris by measuring the Normalized Discounted Cumulative Gain (NDCG) between its cross-modal and intra-modal retrieval results.

Abstract

Image-text alignment relies on large-scale, well-aligned image-text pairs, yet web-crawled datasets inevitably introduce noisy correspondence. Existing methods typically re-weight training pairs by estimating semantic matching degrees from the model's own similarity predictions, which are inherently unreliable under noise and often overestimate partially aligned pairs. To overcome this limitation, we introduce a novel semantic matching degree estimation method based on ranking consistency. Our basic idea is that for a well-aligned image-text pair (I, T), the image I and text T should occupy semantically consistent positions in the embedding space. Thus, using I and T as queries to retrieve other texts (images) should yield two highly consistent rankings. Therefore, we reweight the training paris by measuring the Normalized Discounted Cumulative Gain (NDCG) between its cross-modal and intra-modal retrieval results. This signal is derived from the global data manifold structure rather than point-wise predictions, making it more robust to noise. Experiments on multiple cross-modal retrieval benchmarks validate the effectiveness of our method. Code is available at https://github.com/Aliinton/RCR.

Read PDF

Similar papers

Open access Aug 2026

Dynamic Uncertainty Learning with Noisy Correspondence for Text-based Person Retrieval

Dynamic Uncertainty with Noisy Correspondences (DUNC), a novel framework that incorporates two key components that incorporates Cross-modal Evidential Learning (CEL), which models bidirectional alignment uncertainty using a Dirichlet distribution to capture the confidence in image–text similarity, and Dynamic Robust Lo...

Ze-Qun Xie, Chu-Xin Wang, Si-Hang Cai et al. · 0 citations
Open access 2026

Semantic-Aware Consistency Rectification for Cross-Modal Person Retrieval

Text-to-Image Person Retrieval (TIPR) faces significant challenges due to the “one-to-many” nature of cross-modal matching, where a single identity corresponds to multiple images with varying viewpoints. Standard metric learning approaches often enforce rigid alignment between a text description and all same-identity i...

Zi-Xuan Zhang, Lu-Ming Xiao · 0 citations

CroQS: Cross-modal Query Suggestion for Text-to-Image Retrieval in Image Archives

The novel task of cross-modal query suggestion is introduced, which interactively guides users by suggesting textual refinements based on visual clusters identified in the retrieval results, and the creation of CroQS, a benchmark dataset comprising 50 diverse queries and 295 semantic clusters in generic domain.

Giacomo Pacini, Nicola Messina, Nicola Tonellotto et al. · 0 citations
Conference Open access Sep 2026

Auxiliary text-guided image restoration for image-text matching

A new ITM framework that improves the model's discriminative performance by focusing on localized core attributes and has a better robustness in handling highly similar hard negatives, which provides a new way to address cross-modal hard sample discrimination.

Kuang-Rong Hao · 0 citations
#machine learning Preprint Sep 2026

Spherical Interpolation for Backward-Compatible Multimodal Representations

Contrastive vision-language models map visual and textual representations into a shared normalized embedding space, making cosine similarity the natural metric for cross-modal retrieval. A practical challenge arises during model upgrades: independently trained models generally produce incompatible representation spaces...

Simone Ricci, Niccoló Biondi, F. Pernici · 0 citations
Aug 2026

CATSANet: cross-modal semantic token selection and optimal transport based part alignment for text-to-image person ReID

The proposed Cross-Modal Adaptive Token Selection and Alignment Network (CATSANet), a CLIP-based framework tailored for TI-ReID, achieves competitive performance in terms of Rank-k accuracy and mAP, demonstrating the effectiveness of fine-grained alignment and ranking refinement across datasets.

Dongbin Chen, Junjie Li, Hao Xu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.