Jul 2026· Annual International ACM SIGIR Conference on Research and Development in Information Retrieval· pp. 4211-4215· 0 citations· 28 references
Computer Science
TL;DR
This work aims to explore how well vector-based pseudo relevance feedback can adapt to dense retrieval models when it is not feasible to use surface-form information to pick discriminating expansion tokens.
Abstract
Prior work has shown that vector-based pseudo relevance feedback (PRF) is an effective technique for query expansion for improving retrieval results in dense information retrieval. In dense retrieval, ColBERT-PRF has emerged as a novel mechanism, using cluster centroids built from feedback documents as PRF expansion tokens and leveraging statistical information from the closest neighboring token ids to dictate how useful these expansion tokens are. While this approach has been shown to work well in the monolingual retrieval setting for English using the original ColBERT infrastructure, such systems have since evolved to improve inference speed, reduce storage and memory usage, and support cross-language (CLIR) and multilingual (MLIR) retrieval. As a result, many of these advancements have reduced the ability to utilize token-level statistics. In this work, we aim to explore how well this type of approach can adapt to dense retrieval models when it is not feasible to use surface-form information to pick discriminating expansion tokens. Furthermore, we explore alternative clustering mechanisms, such as HDBScan, to compare how different clustering methods perform at building clusters that can be useful for PRF. Experiments on MLIR, CLIR, and Report Generation tasks, such as those in the TREC 2024 NeuCLIR Report Generation Pilot Task, show that even without access to these token statistics, the use of cluster centroids for PRF can still improve nDCG and α-nDCG by up to 12%.
This work addresses the problem of ranking dimensions of dense vectors by proposing and evaluating methods coming from different conceptual backgrounds and shows that the first dimensions contribute the most to total effectiveness performance when ranking the dimensions with the best-performing methods.
Gabriel H. Tolosa, Tomás Delvechio, E. A. Ríssola· International Journal of Dat...· 0 citations
This work introduces AnchorQE, a training-free method that separately encodes the original query and its expansion before interpolating them, and shows that AnchorQE improves retrieval effectiveness by up to 12.89% when compared to widely-used expansion-only or text-level concatenation baselines across TREC-DL, LoTTE,...
A systematic comparison of retrieval strategies for candidate generation under a shared LLM-based selection stage, combining sparse retrieval (BM25), Web KB search, and a state-of-the-art trained dense retriever with several open- and closed-source LLMs is presented.
Fina Polat, Daniel Daza, Pengyu Zhang et al.· 0 citations
This paper examines three major families of retrieval models that have shaped this evolution of information retrieval: vector space models, probabilistic retrieval, and neural retrieval, and shows that modern search systems rarely rely on a single model.
Prapitha Gopi K· International Journal of Tec...· 0 citations
It is suggested that embedding-based retrieval is more effective in capturing semantic information, including synonym usage and implicit contextual relationships, within the evaluated dataset, and compact embedding models such as MiniLM may provide an alternative approach to traditional keyword-based retrieval methods...
Ilham Yusuf Faturochman, Aprilisa Arum Sari, Nibras Faiq Muhammad· IC-ITECHS· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.