Skip to content
Preprint

Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval

Aug 2026 · 0 citations · 41 references
Computer Science

TL;DR

This work introduces UniME-R1, an embedder-adviser framework that learns to reason over initially retrieved candidates and generate Retrieval-Centric Chain-of-Thought (RC-CoT), and aligns the adviser with retrieval outcomes through supervised learning and retrieval-oriented reinforcement learning.

Abstract

Unified multimodal retrieval aims to identify candidates that satisfy complex user intent expressed through heterogeneous inputs. Although Large Vision-Language Model (LVLM)-based retrievers are efficient and scalable, directly encoding raw multimodal inputs often misses fine-grained discriminative cues, leading to confusion among semantically similar candidates. Recent methods mitigate this limitation by generating Chain-of-Thought (CoT) rationales to enrich the query representation. However, such reasoning is typically derived from the query alone: it explains what the query describes, but not what the retriever misunderstands. We argue that effective retrieval reasoning should instead be conditioned on retrieval feedback. Based on this insight, we introduce UniME-R1, an embedder-adviser framework that learns to reason over initially retrieved candidates and generate Retrieval-Centric Chain-of-Thought (RC-CoT). The adviser analyzes candidates individually to identify the discriminative cues confused by the embedder. If the target appears in the initial top-k set, UniME-R1 directly reranks the candidates; otherwise, it generates RC-CoT to refine the retrieval direction and performs full-corpus re-retrieval with a dual-mode embedder. To train the framework, we mine hard negatives to simulate realistic retrieval failures, jointly optimize direct retrieval and RC-CoT-augmented retrieval, and align the adviser with retrieval outcomes through supervised learning and retrieval-oriented reinforcement learning. Extensive experiments on MMEB-V2 and a diverse set of general multimodal retrieval benchmarks demonstrate that UniME-R1 consistently improves retrieval performance over strong baselines.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

UMER: Unifying Embedding and Ranking via Pair-Aware Discriminative Reasoning for Universal Multimodal Retrieval

UMER replaces item-wise reflection with Pair-Aware Discriminative Reasoning, which compares query--candidate pairs to identify instruction-relevant matching and discrepancy evidence and achieves state-of-the-art performance under comparable experimental settings while supporting budget-adjustable inference.

Libiao Chen, Xiyang Liu, Yanheng Wei et al. · 1 citation
#artificial intelligence Preprint Sep 2026

Reason What Matters: Retrieval-Grounded Reasoning for Universal Multimodal Embeddings

Universal multimodal embedding (UME) learns unified representations across modalities, enabling a single model to support diverse retrieval tasks. Recent methods use Chain-of-Thought (CoT) reasoning to better interpret multimodal inputs before generating embeddings for complex retrieval tasks and further optimize this...

Mingzhou Jiang, Pei-Xi Wu, Hang Cheng et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Beyond One-Shot Expansion: Contrastive Evidence Exploration for Multi-Hop Retrieval

This work proposes a training-free multi-hop retrieval framework that integrates evidence-conditioned exploration, passage-specific contrastive refinement, and coverage-aware final ranking and demonstrates consistent improvements in retrieval quality and downstream QA performance over baselines.

Jungmin Yun, Youngbin Kim · 0 citations
Book Open access Aug 2026

Retrv-MoE: Scaling Unified Multimodal Retrieval with Sparse Mixture-of-Experts

This work proposes Retrv-MoE, a unified retrieval architecture built upon sparse Mixture-of-Experts (MoE), and theoretically and empirically demonstrates that this conditional computation mechanism provides a structural remedy to optimization interference by decoupling the learning trajectories of conflicting tasks and...

Tongxu Lin, Jiayin Xiao · 0 citations
Preprint Aug 2026

UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering

UniHEAR is a unified lightweight framework for heterogeneous-source entity retrieval and reranking that achieves state-of-the-art retrieval and VQA performance, improving Recall@1 and Recall@1 by 6.7 and 1.2 points over the strongest baselines while maintaining a lightweight reranking architecture.

Ganzhong Luo, Yang Ren, Han-Yong Wang et al. · 0 citations
Preprint Aug 2026

Beyond Visual Similarity: Entity-Aligned Retrieval for Knowledge-Based Visual Question Answering

KBMR is proposed, the first MLLM-based embedding retriever tailored for KB-VQA, and an MLLM-based semantic discriminator that generates continuous entity-consistency weights is introduced to tackle the challenge of noisy supervision in Wikipedia-scale retrieval.

Hangrui Xu, Zheng-Xian Wu, Yu Yu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.