Skip to content

Effective query-independent vector pruning for dense retrieval

Jul 2026 · International Journal of Data Science and Analysis · Vol 22 · 0 citations · 45 references

TL;DR

This work addresses the problem of ranking dimensions of dense vectors by proposing and evaluating methods coming from different conceptual backgrounds and shows that the first dimensions contribute the most to total effectiveness performance when ranking the dimensions with the best-performing methods.

View source

Similar papers

Preprint Aug 2026

AdaWidth: Query-Adaptive Embedding Width for Dense Retrieval

AdaWidth is introduced, which adapts the number of evaluated dimensions to each query within a shared prefix representation, and derives a prefix sufficiency analysis showing that the required number of dimensions is set by the competing documents at the retrieval cutoff.

Shu-Bing Yang, Dongfang Zhao · 0 citations
Preprint Aug 2026

Retrieval Needs Multivectors: An Exponential Separation

This work provides the first explicit family of query and document sets, together with their relevance matrices, for which single-vector embeddings that rank all relevant documents above irrelevant ones require exponential size, whereas polynomial-size multi-vector embeddings suffice.

Mihir Agarwal, Viraj Agrawal, Sabyasachi Basu et al. · 1 citation
#small language model Preprint Aug 2026

Query Expansion Is More Than Generation: Improving Dense Retrieval through Better Integration

This work introduces AnchorQE, a training-free method that separately encodes the original query and its expansion before interpolating them, and shows that AnchorQE improves retrieval effectiveness by up to 12.89% when compared to widely-used expansion-only or text-level concatenation baselines across TREC-DL, LoTTE,...

Sixia Sun, Mihai Surdeanu · 0 citations
Preprint Sep 2026

Pre-retrieval Query Clustering for Adaptive Top-k Document Retrieval in RAG Systems

RAG systems commonly retrieve a fixed number of documents (top-k) to ground generation, but this static approach is brittle: simple queries suffer over-retrieval (adding noise and cost) while complex queries are under-retrieved, causing recall failures that cascade into incorrect answers. Motivated by the question of h...

Ye Xia, Emre Yamangil, Hai-Xun Wang · 0 citations
Review Open access Aug 2026

From Vector Space to Neural Ranking: A Comparative Study of Modern Information Retrieval Models

This paper examines three major families of retrieval models that have shaped this evolution of information retrieval: vector space models, probabilistic retrieval, and neural retrieval, and shows that modern search systems rarely rely on a single model.

Prapitha Gopi K · 0 citations
Preprint Aug 2026

Hypergraph Embedding Indexing for Efficient Dense Vector Retrieval

Dense vector retrieval has become the foundation of modern semantic search, yet existing approximate nearest neighbor (ANN) indexes treat an embedding as an indivisible point in a high-dimensional space. In this work, we propose the Hypergraph Embedding Index (HEI), a framework that instead organizes documents accordin...

Kishore Konda · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.