Using three support-annotated multi-hop QA benchmarks, this work compares matched adapted readers trained with raw context, retrieval windows, and gold-support diagnostic renderings to distinguish support-availability failures from remaining reader-interface effects.
Abstract
In multi-hop RAG evaluation, a top-k answer score can hide two different failures: the retrieval window may drop part of the support chain, or it may contain support in a form the adapted reader does not use well. We call this reader-facing form of retrieved evidence an evidence interface. Using three support-annotated multi-hop QA benchmarks, we compare matched adapted readers trained with raw context, retrieval windows, and gold-support diagnostic renderings. These comparisons distinguish support-availability failures from remaining reader-interface effects. Top-k windows become interpretable only after checking whether the complete annotated support chain survives: when it does, short ranked windows can match or improve over raw context; when it does not, missing support explains much of the loss. Gold support-first improves matched readers; on 2Wiki and MuSiQue, a support-supervised ranker raises coverage and recovers raw-context quality at lower prompt cost, while retaining gold headroom. Support-removal checks further show that the gains rely on exposed evidence, not only answer priors. On support-annotated evaluations, top-k answer scores should therefore be reported together with complete-support coverage.
Retrieval-augmented generation (RAG) improves large language models by grounding generation in external evidence, but it also introduces a source trust problem: retrieved context may be useful, irrelevant, or even misleading. Existing RAG systems often apply a fixed trust policy toward retrieved evidence, which can eit...
Retrieval evaluation for retrieval-augmented generation (RAG) is increasingly designed around whether retrieved passages contain evidence that can support generation, rather than topical relevance alone. We study whether this closer alignment with downstream evidence needs also makes retrieval evaluation more useful fo...
Utshab Kumar Ghosh, D. Mukhopadhyay, Shubham Chatterjee· 0 citations
The described Virtual Human demonstration system uses RAG with 14 geriatrics patient brochures from the Canisius Wilhelmina Hospital, which were converted into question-answer chunks and embedded and retrieved based on similarity to user queries.
Roel Boumans· Proceedings of the 26th ACM...· 0 citations
MCoRe, a multi-entry complementary retrieval framework with reflection-guided iteration for multi-hop QA that enables multi-entry complementary retrieval by indexing entry units at multiple semantic resolutions with explicit links to chunk evidence, and fusing cross-resolution hits via chunk-level voting to form a comp...
Ju-Xiang Zeng, Zhuohui Gao, Zhe Hou et al.· Proceedings of the 32nd ACM...· 0 citations
EviRank is recast multimodal image re-ranking as a semantic constraint satisfaction problem and proposed, which parses any query - text-only, image-only, or composed - into a unified evidence package: typed criteria across six semantic slots, each labelled required, forbidden, or ignorable.
Enjun Du, Siyi Liu, Zi-Rong Chen et al.· 2 citations
MIDR (Multimodal Indexing for Document Retrieval), a training-free framework for enrichment-augmented indexing that shifts multimodal reasoning to index time, establishes index-time multimodal reasoning as a compelling accuracy-deployment alternative to serving-time visual late interaction.
Debanjan Mahata, Atharva Tendle, Daniel Preoţiuc-Pietro et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.