Skip to content

TurboVec: A Case Study in Cost-Efficient Private Retrieval for Enterprise RAG via Codebook-Oblivious Quantization

Jul 2026 · arXiv.org · Vol abs/2607.16973 · 1 citation · 26 references
Computer Science

TL;DR

TurboVec, an open-source vector index built on TurboQuant - a codebook-oblivious scalar quantizer requiring no corpus-dependent training is studied, which achieves 11ms median query latency at 100K vectors versus 707ms for warehouse brute-force scan.

Abstract

Retrieval-Augmented Generation (RAG) systems increasingly power enterprise LLM applications, yet the vector retrieval layer introduces two underexplored challenges: (1) trained codebook quantizers may expose corpus statistics during index construction, creating a leakage channel in multi-tenant deployments, and (2) post-hoc filtering for tenant isolation degrades recall on selective queries. We study TurboVec, an open-source vector index built on TurboQuant - a codebook-oblivious scalar quantizer requiring no corpus-dependent training. On the DBpedia OpenAI embeddings benchmark (d=1536, 100K-999K vectors), TurboQuant 4-bit outperforms trained FAISS Product Quantization at the same memory budget by 8.5-8.9 percentage points in Recall@5 across all scales. Compared to HNSW (R@5=0.991) and IVF-PQ (R@5=0.840), TurboQuant occupies a distinct design point: higher recall than IVF-PQ without training, at 4-8x less memory than HNSW. Deployed on Snowpark Container Services, TurboVec achieves 11ms median query latency at 100K vectors versus 707ms for warehouse brute-force scan. Kernel-level allowlist filtering maintains 0.86-0.93 Recall@10 across 10-1000 tenant workloads versus 0.09-0.19 for post-filter baselines. Codebook-oblivious design reduces membership inference accuracy to near-random (50.0%) versus 57.3% for PQ codebooks. Limitations include single dataset evaluation, uncompressed HNSW comparison, and privacy evaluation on synthetic data only.

View source

Similar papers

#artificial intelligence Preprint Oct 2026

MRVQ: One Resident Index for Dimension- and Rate-Elastic Vector Search

Dense-retrieval services must switch among embedding-prefix dimensions and index bit rates as latency, quality, and memory budgets change. Tuning a quantizer separately for each rate gives the best quality, but the retrieval tier then holds several code streams and quantizer states at once. We introduce Matryoshka Resi...

S. Culatana, Shan Huang, Kang Li · 0 citations
#artificial intelligence Preprint Sep 2026

Matryoshka Hash Representations for Model-Aware Compact Semantic Retrieval

This work introduces Matryoshka Hash Representations (MHR), a two-stage procedure that separates full-width training from prefix organization and strengthens two common pipelines: shortlisting candidates for full-precision reranking, and pruning a low-storage graph index such as LEANN.

Pei-Chun Hua, Yun-Ming Xiao · 1 citation
#artificial intelligence Preprint Sep 2026

All for 1-Bit: Towards Genuine 1-Bit Post-Training Quantization for LLMs

All for 1-Bit (AF1) is proposed, a genuine 1-bit PTQ framework for LLMs that consistently outperforms existing binarization-based PTQ methods in perplexity and zero-shot accuracy, providing a practical path toward deployable genuine 1-bit compression for LLMs.

Zhi-Xiong Zhao, Zu-Kang Xu, Guang-Yu Sun et al. · 1 citation
#machine learning Preprint Sep 2026

Spruce: Scalable Private Outsourced Retrieval Using Compact Embeddings

Spruce learns compact binary codes that preserve candidates for full-precision reranking, replacing corpus-wide embedding scoring with efficient Hamming-distance computation under two-server multi-party computation under two-server multi-party computation (MPC).

Pei-Chun Hua, Yun-Ming Xiao · 3 citations
Review Aug 2026

Transforms for LLM Quantization: The Great Inversion and Format Co-Design

This work identifies and formalizes the principle that organizes the Great Inversion, the Great Inversion: allocation-flexible coding rewards energy concentration, whereas the grouped shared-scale quantization a deployed matrix instruction performs rewards within-group flattening.

Ehsan Jokar · 0 citations
#artificial intelligence Preprint Oct 2026

Tailoring the Quantization Space for 1-Bit KV Cache Compression

The key-value (KV) cache becomes a major memory bottleneck in long-context LLM inference, placing substantial pressure on memory capacity and bandwidth. To mitigate this bottleneck, vector quantization (VQ) has emerged as a promising approach for aggressive KV cache compression. However, existing VQ methods degrade sub...

Minsoo Cheong, Donghyun Son, Sungjoo Yoo · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.