Skip to content
Open access

Automated Caption-Guided Image Retrieval: A Multi-Positive Contrastive Learning Paradigm

Aug 2026 · Frontiers in Science and Engineering · 0 citations · 22 references

TL;DR

The proposed Multi-Positive Example Contrastive Learning (MPC) model provides more discriminative sentence embeddings and improves semantic similarity measurement in image retrieval.

Abstract

[Objective] This study aims to improve textual similarity measurement for image retrieval while mitigating the anisotropy of sentence embeddings. [Methods] We propose an image-caption semantic encoder trained with a multi-positive contrastive loss. The conventional contrastive objective is extended to accommodate multiple positive samples, and image captioning is used to generate training data automatically. On this basis, we construct a caption-based image-to-image retrieval framework. [Results] Experiments show that the proposed model outperforms baseline methods on semantic textual similarity (STS) benchmarks and improves the agreement between retrieval results and human semantic judgments. [Limitations] Short captions cannot fully represent the complex semantics, ambiguity, and fine-grained details of an image. [Conclusion] The proposed Multi-Positive Example Contrastive Learning (MPC) model provides more discriminative sentence embeddings and improves semantic similarity measurement in image retrieval.

Read PDF

Similar papers

Open access Jul 2026

A Unified Multimodal Search Framework Using Generative AI and Image Understanding for Enhanced Information Retrieval

A unified multimodal search framework that integrates generative AI-based captioning and image understanding for improved retrieval, enabling more accurate, context-aware search in applications such as e-commerce, multimedia, and large-scale retrieval.

Saeed Alzahrani, Farah Mohammad, Nazar Hussain · 0 citations
Conference Jul 2026

Multimodal Context-Enriched Visual Representation Learning for Enhanced Vision–Language Image Captioning

Image captioning models can produce rapid Sentences, without visual relationships, or insert non-existing plausible objects. A common cause is to compress image evidence into visual symbols that carry a weak neighbourhood context. The multimodal context-enhanced visual representation learning framework (MCVRL) addresse...

E. Divya, Johnson Kolluri, Kiran Siripuri · 0 citations
Preprint Aug 2026

CoCo-IR: Contextual Composed Image Retrieval

A new model based on a Large Multimodal Model (LMM) that functions as a context-aware reasoner for CoCo-IR is proposed, which interprets the entire interaction history to generate Transformable Image Embeddings (TIE) that evolve across turns.

Shengcao Cao, T. Dabral, Z. Ding et al. · 0 citations
Review Aug 2026

Structuring Semantic Embeddings for Principle Evaluation: A Prototype-Guided Contrastive Learning Approach

This paper introduces Prototype-Guided Contrastive Learning (PGCL), a prototype-guided geometric regularization module built on top of frozen text embeddings, and introduces Prototype-Guided Contrastive Learning (PGCL), a prototype-guided geometric regularization module built on top of frozen text embeddings.

Che Shen, Junwei Su, Ling-Peng Kong et al. · 1 citation
Preprint Aug 2026

Rethinking Text-Based Image Retrieval in Specific Domain

The Semantic-Aware Fine-Tuning (SAFT) framework is proposed to address semantic compression in specific domains, which incorporates Semantic-Aware Soft-Label Supervision and Intra-modal Structural Distillation to establish a promising paradigm for domain-specific TBIR tasks.

Jing-Yang Tan, Shengan Yang, Yuanpeng Chen et al. · 0 citations
Jul 2026

Probabilistic Embeddings With Evidence Learning and Refinement for Text–Video Retrieval

This paper studies the problem of text-video retrieval, where the goal is to learn accurate cross-modal alignment between videos and text. This problem is challenging because of the matching ambiguity caused by the inherent gap between the heterogeneous video and text modalities. In particular, the differences in the i...

Dong-Lin Zhang, Zheng-Hao Rao, Xing Xu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.