Jul 2026· International Conference on the Theory of Information Retrieval· pp. 73-81· 0 citations· 77 references
Computer Science
TL;DR
This work presents a study of ranking approaches, including lexical, dense and zero-shot prompted large language models (LLMs), to rank passages in relevant documents based on the presumed fraction of relevant text they contain, and demonstrates the merits of these approaches in utilizing relevance feedback.
Abstract
The ad hoc document retrieval task is to rank documents by their presumed relevance to a query. Most TREC benchmarks provide relevance judgments only at the document level, without indicating which parts of a document are actually relevant. Focused relevance judgments, which highlight relevant text at the character level, are valuable but scarce. In this work, we present a study of ranking approaches, including lexical, dense and zero-shot prompted large language models (LLMs), to rank passages in relevant documents based on the presumed fraction of relevant text they contain. Our analysis shows that LLM-based rankings are highly effective and outperform strong sparse and dense retrieval baselines. We demonstrate the merits of our approaches in utilizing relevance feedback: constructing relevance models from top-ranked passages in relevant documents yields performance that transcends that of relevance models constructed from the entire documents.
This study empirically evaluates the robustness of an IR model to the addition of non-relevant documents by merging two collections with negligible topic overlap and finds that MDA is more effective than MDD for retrieval, whereas MDD and MDA rerankers are equally effective.
Emmanouil Georgios Lionis, Sean MacAvaney, Debasis Ganguly· 0 citations
It is concluded that knowing the relevance of entities for an information need can be very valuable, but that the methods described are insufficient to determine this relevance information from the documents and queries alone.
Norbert Boudens, Chris Kamphuis, A. D. de Vries· Annual International ACM SIG...· 1 citation
Evaluation on WikiSA and ExaRank shows that ranking-based few-shot prompting generally improves over zero-shot prompting and achieves competitive performance against random-shot prompting, indicating that retrieval-based demonstration selection is beneficial but not uniformly superior in all settings.
A. Laksito, Aali Alqarni, Mark Stevenson· International Conference on...· 0 citations
It is found that focused retrieval techniques clearly outperform document retrieval, especially for the Focused and Restricted Relevant in Context Tasks, which limit the amount of text than can be returned per topic and per article respectively.
It is shown that duplicate redundancy and LLM paraphrasing does not significantly improve answer correctness, however, providing diverse documents is highly beneficial, improving answer correctness by 17%-47%.
Jonathan J. Ross, B. Koopman, A. H. van der Vegt et al.· 2 citations