Sep 2026· Proceedings of the Thirty-Fifth International Joint Conference on Artificial Intelligence· pp. 4705-4713· 0 citations· 38 references
TL;DR
A novel semantic matching degree estimation method based on ranking consistency that reweights the training paris by measuring the Normalized Discounted Cumulative Gain (NDCG) between its cross-modal and intra-modal retrieval results.
Abstract
Image-text alignment relies on large-scale, well-aligned image-text pairs, yet web-crawled datasets inevitably introduce noisy correspondence. Existing methods typically re-weight training pairs by estimating semantic matching degrees from the model's own similarity predictions, which are inherently unreliable under noise and often overestimate partially aligned pairs. To overcome this limitation, we introduce a novel semantic matching degree estimation method based on ranking consistency. Our basic idea is that for a well-aligned image-text pair (I, T), the image I and text T should occupy semantically consistent positions in the embedding space. Thus, using I and T as queries to retrieve other texts (images) should yield two highly consistent rankings. Therefore, we reweight the training paris by measuring the Normalized Discounted Cumulative Gain (NDCG) between its cross-modal and intra-modal retrieval results. This signal is derived from the global data manifold structure rather than point-wise predictions, making it more robust to noise. Experiments on multiple cross-modal retrieval benchmarks validate the effectiveness of our method. Code is available at https://github.com/Aliinton/RCR.
Dynamic Uncertainty with Noisy Correspondences (DUNC), a novel framework that incorporates two key components that incorporates Cross-modal Evidential Learning (CEL), which models bidirectional alignment uncertainty using a Dirichlet distribution to capture the confidence in image–text similarity, and Dynamic Robust Lo...
Ze-Qun Xie, Chu-Xin Wang, Si-Hang Cai et al.· ACM Transactions on Informat...· 0 citations
Text-to-Image Person Retrieval (TIPR) faces significant challenges due to the “one-to-many” nature of cross-modal matching, where a single identity corresponds to multiple images with varying viewpoints. Standard metric learning approaches often enforce rigid alignment between a text description and all same-identity i...
The novel task of cross-modal query suggestion is introduced, which interactively guides users by suggesting textual refinements based on visual clusters identified in the retrieval results, and the creation of CroQS, a benchmark dataset comprising 50 diverse queries and 295 semantic clusters in generic domain.
Giacomo Pacini, Nicola Messina, Nicola Tonellotto et al.· 0 citations
A new ITM framework that improves the model's discriminative performance by focusing on localized core attributes and has a better robustness in handling highly similar hard negatives, which provides a new way to address cross-modal hard sample discrimination.
Kuang-Rong Hao· International Conference on...· 0 citations
Contrastive vision-language models map visual and textual representations into a shared normalized embedding space, making cosine similarity the natural metric for cross-modal retrieval. A practical challenge arises during model upgrades: independently trained models generally produce incompatible representation spaces...
Simone Ricci, Niccoló Biondi, F. Pernici· 0 citations
The proposed Cross-Modal Adaptive Token Selection and Alignment Network (CATSANet), a CLIP-based framework tailored for TI-ReID, achieves competitive performance in terms of Rank-k accuracy and mAP, demonstrating the effectiveness of fine-grained alignment and ranking refinement across datasets.
Dongbin Chen, Junjie Li, Hao Xu et al.· Pattern Analysis and Applica...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.