Jul 2026· International Journal of Data Science and Analysis· Vol 22· 0 citations· 45 references
TL;DR
This work addresses the problem of ranking dimensions of dense vectors by proposing and evaluating methods coming from different conceptual backgrounds and shows that the first dimensions contribute the most to total effectiveness performance when ranking the dimensions with the best-performing methods.
AdaWidth is introduced, which adapts the number of evaluated dimensions to each query within a shared prefix representation, and derives a prefix sufficiency analysis showing that the required number of dimensions is set by the competing documents at the retrieval cutoff.
This work provides the first explicit family of query and document sets, together with their relevance matrices, for which single-vector embeddings that rank all relevant documents above irrelevant ones require exponential size, whereas polynomial-size multi-vector embeddings suffice.
Mihir Agarwal, Viraj Agrawal, Sabyasachi Basu et al.· 1 citation
This work introduces AnchorQE, a training-free method that separately encodes the original query and its expansion before interpolating them, and shows that AnchorQE improves retrieval effectiveness by up to 12.89% when compared to widely-used expansion-only or text-level concatenation baselines across TREC-DL, LoTTE,...
RAG systems commonly retrieve a fixed number of documents (top-k) to ground generation, but this static approach is brittle: simple queries suffer over-retrieval (adding noise and cost) while complex queries are under-retrieved, causing recall failures that cascade into incorrect answers. Motivated by the question of h...
This paper examines three major families of retrieval models that have shaped this evolution of information retrieval: vector space models, probabilistic retrieval, and neural retrieval, and shows that modern search systems rarely rely on a single model.
Prapitha Gopi K· International Journal of Tec...· 0 citations
Dense vector retrieval has become the foundation of modern semantic search, yet existing approximate nearest neighbor (ANN) indexes treat an embedding as an indivisible point in a high-dimensional space. In this work, we propose the Hypergraph Embedding Index (HEI), a framework that instead organizes documents accordin...
Kishore Konda· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.