Skip to content

EigenLI: Spectral Approximations to Late Interaction

Sep 2026 · 0 citations · 49 references
Computer Science

TL;DR

This work introduces EigenLI, a spectral approximation framework that compresses late-interaction representations via document-specific low-dimensional subspaces and consistently outperforms comparable single-vector surrogates such as MUVERA.

Abstract

Late-interaction models such as ColBERT achieve strong effectiveness by representing each document with many token-level vectors, but this expressivity leads to large indexing cost, storage footprints and expensive MaxSim scoring. We show that late-interaction representations exhibit an intrinsic low-rank structure: document token embeddings concentrate in a low-dimensional subspace that preserves most of the retrieval signal. Leveraging this observation, we introduce EigenLI, a spectral approximation framework that compresses late-interaction representations via document-specific low-dimensional subspaces. Unlike clustering or pooling methods, EigenLI identifies the dominant eigendirections of each document and uses them to construct reduced interaction representations. Empirically, $k$-EigenLI with $k \le 32$ outperforms k-means and Ward clustering based pooling methods on ColBERTv2 and AnswerAI-ColBERT-small; GTE-ModernColBERT exhibits a different tradeoff at $k=32$, where clustering methods perform better. The same spectral construction also yields EigenLI-SV, an ANN-compatible single-vector representation derived from the second-order summary of the reduced structure. Across multiple datasets and all three text models, EigenLI-SV consistently outperforms comparable single-vector surrogates such as MUVERA.

View source

Similar papers

Preprint Sep 2026

Generative Late-Interaction Embeddings For Visual Document Retrieval

Late-interaction retrieval is the state-of-the-art for visual document search, but it pays for its accuracy in storage. Existing compression methods retain a subset or local average of the N~1,000 vectors per page. Under aggressive storage budgets, however, these methods degrade sharply, and alternatives require retrai...

M. Eltahir, Talal Aloushan, Rose Khairoalsendi et al. · 0 citations
Preprint Aug 2026

Retrieval Needs Multivectors: An Exponential Separation

This work provides the first explicit family of query and document sets, together with their relevance matrices, for which single-vector embeddings that rank all relevant documents above irrelevant ones require exponential size, whereas polynomial-size multi-vector embeddings suffice.

Mihir Agarwal, Viraj Agrawal, Sabyasachi Basu et al. · 1 citation
#artificial intelligence Preprint Aug 2026

A Manifold-Aware Topic Modeling Approach via Rank-Based Prototypes

Recent topic models leverage pretrained embeddings, but neural architectures produce latent representations without grounding in specific texts, and clustering-based pipelines assign representative documents only post hoc, relying on absolute distances distorted by hubness and anisotropy in high-dimensional spaces. We...

Thiago César Castilho Almeida, D. Pedronette · 0 citations
#artificial intelligence Preprint Aug 2026

Spatial Matryoshka Training for Multi-Granularity Visual Document Retrieval

It is demonstrated that models trained using ColSNAP maintain near full-resolution retrieval performance under substantial compression and that ColSNAP transfers effectively across multiple late-interaction backbones, and achieves most of its improvements via a lightweight adaptation stage applied to a pre-trained retr...

Trishan Singha Roy, Arkadeep Acharya, Vishwajeet Kumar et al. · 0 citations
Preprint Sep 2026

AdaMerge: Tuning-Free Patch Compression for Multi-Vector Visual Document Retrieval

Multi-vector visual document retrieval (VDR) models such as ColPali and ColNomic achieve strong accuracy by representing each document with hundreds to thousands of patch-level embeddings, at substantial storage and latency cost. Existing compression methods either prune unimportant patches or merge similar ones into c...

Jian-Xin You, Kun Ni · 0 citations
Preprint Aug 2026

Coverage Matters: MarginMerge for Compressing Multi-Vector Visual Document Retrievers

It is argued that effective compression should preserve query-relevant coverage, meaning the diverse document regions that may become the strongest MaxSim match across queries, rather than selecting patches independently by salience, why dense rendered pages are easier to compress than natural images.

Ailar Mahdizadeh, Aria Salari, Sohail Rajabi et al. · 0 citations

Related blog posts

Microsoft Research Blog Sep 30, 2026

Forecasting space weather risks on power grids

Extreme space-weather events can damage power systems on Earth and degrade GPS accuracy and satellite operations. A new machine learning system can predict where damage is likely to occur 30-60 minutes before a storm arrives. The post Forecasting space weather risks on power grids appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.