Skip to content
Preprint

Retrieval Needs Multivectors: An Exponential Separation

Aug 2026 · 1 citation · 28 references
Computer Science

TL;DR

This work provides the first explicit family of query and document sets, together with their relevance matrices, for which single-vector embeddings that rank all relevant documents above irrelevant ones require exponential size, whereas polynomial-size multi-vector embeddings suffice.

Abstract

Recent works have highlighted the expressive limitations of embedding based retrieval models through both theoretical analyses and challenging benchmarks such as LIMIT. While multi-vector embeddings consistently outperform single-vector embeddings, the precise representational gap between them remains poorly understood. In this work, following Jayaram's work, we provide the first explicit family of query and document sets, together with their relevance matrices, for which single-vector embeddings that rank all relevant documents above irrelevant ones require exponential size, whereas polynomial-size multi-vector embeddings suffice. Our result establishes an exponential separation between the expressive power of single-vector and multi-vector embeddings for the task of ranking of documents as opposed to approximating numerical scores as in the work of Jayaram. Motivated by our theoretical construction, we introduce ANDOR, a new retrieval benchmark that naturally instantiates these hard examples. We show that state-of-the-art single-vector embedding models perform poorly on ANDOR in the zero-shot setting and exhibit only marginal improvements after fine-tuning, highlighting the inherent difficulty of the benchmark compared to prior work. In contrast, multi-vector models consistently outperform their single-vector counterparts and improve substantially with fine-tuning, closely aligning with our theoretical predictions.

View source

Similar papers

Preprint Aug 2026

AdaWidth: Query-Adaptive Embedding Width for Dense Retrieval

High-dimensional embeddings are central to dense retrieval, but not all of these dimensions need to be evaluated at retrieval time. Existing methods reduce dimensions in two ways: truncating the same leading dimensions for every query, or masking a different subset for each query while still storing and accessing the f...

Shu-Bing Yang, Dongfang Zhao · 0 citations
Preprint Aug 2026

Giga-Embeddings: Mixture-of-Experts Encoders for High-Throughput Text Embeddings

This work introduces Giga-Embeddings, a family of text embedding models designed to combine strong retrieval quality with efficient serving, and trains the compact model using a dimension-agnostic objective that aligns teacher and student similarity distributions.

Egor Kolodin, Egor Krasnoperov, Evgeniy Kosarev et al. · 0 citations
Preprint Sep 2026

Generative Late-Interaction Embeddings For Visual Document Retrieval

Late-interaction retrieval is the state-of-the-art for visual document search, but it pays for its accuracy in storage. Existing compression methods retain a subset or local average of the N~1,000 vectors per page. Under aggressive storage budgets, however, these methods degrade sharply, and alternatives require retrai...

M. Eltahir, Talal Aloushan, Rose Khairoalsendi et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Shadow Queries for Private Retrieval in Vector Databases

SHAQ (shadow query generation), a semantic-decomposition and embedding-decoupling defense against EIAs, is proposed, which uses a generative language model to create diverse shadow queries that capture different semantic aspects of each document.

Xinguo Feng, Zhongkui Ma, Zi-Han Wang et al. · 0 citations
Preprint Aug 2026

Learning Sample-wise Rank-aware Interpolation Weights for Composed Visual Data Retrieval

This work revisits the efficacy of simple linear interpolation within an embedding space, and introduces SRAIN, the first framework that dynamically predicts query-specific interpolation weights, and achieves the best in composed video retrieval and matches the current state of the art in composed image retrieval.

Boseung Jeong, T. Park, Donghyeon Kwon et al. · 1 citation
Open access Aug 2026

Do Retrieval-Trained Embeddings Help Linear Contextual Bandits?

Text embeddings from retrieval-tuned (dual-encoder) models are increasingly used as context features in contextual bandits for recommendation, on the assumption that an embedding space optimized for inner-product similarity will speed up a linear exploration policy. This study tests that assumption with a controlled, s...

Mustafa Canim · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.