Skip to content
Preprint

ByteX: A Unified AI Search Engine at ByteDance

Aug 2026 · 0 citations · 47 references
Computer Science

TL;DR

ByteX introduces a quantization-aware vector kernel based on SymRaBitQ, a new symmetric quantization scheme with tight theoretical guarantees that allows index construction to run directly in the quantized space accurately and efficiently without retaining a copy of full-precision vectors.

Abstract

Since 2016, ByteX has been the foundation of ByteDance's search infrastructure, scaling to more than 7,000 clusters and 300 PB of indexed data. Driven by the demands of AI workloads, ByteX has evolved from a text search engine into a unified AI search system supporting vector retrieval, lexical matching, and predicate filtering. Its largest deployment indexes nearly one trillion high-dimensional vectors. This scale exposes two central bottlenecks in AI-era retrieval: memory-intensive graph-index construction under sustained ingestion, and the prohibitive cost of keeping vector indexes entirely in memory. ByteX addresses these bottlenecks with two techniques. First, it introduces a quantization-aware vector kernel based on SymRaBitQ, a new symmetric quantization scheme with tight theoretical guarantees that allows index construction to run directly in the quantized space accurately and efficiently without retaining a copy of full-precision vectors. Second, it provides a hybrid storage engine that supports memory-resident, hybrid, and SSD-resident deployments, with fine-grained record-level caching to trade memory for latency under operational control. On large-scale benchmarks, ByteX improves throughput by up to 3x, reduces indexing memory by 80%, and lowers operating cost by 86% compared with prior systems, while supporting trillion-vector scale, write-heavy or latency-sensitive workloads in production.

View source

Similar papers

Preprint Aug 2026

InSituANN: Revisiting IVF for PCIe-Efficient Billion-Scale Vector Search

Approximate nearest neighbor search (ANNS) over billion-scale vector datasets has become a foundational operator for modern retrieval systems, powering large-scale recommendation, semantic search, and LLM/RAG workloads. Although GPUs offer massive parallelism and high-bandwidth memory for batched vector search, their l...

Yue-Meng Xu, Zongxi Liu, Junyu Long et al. · 0 citations
Preprint Sep 2026

Hoss: Fast Oblivious Semantic Search with Heterogeneous GPU-CPU-TEE Architecture

This work proposes Hoss, a first-of-its-kind oblivious semantic search system with a heterogeneous CPU-GPU TEE architecture that supports fast, scalable search with low cost of ownership and features a host-access ORAM mechanism that goes beyond traditional performance constraints and incorporates several data-dependen...

Jian Du, Wei-Jie Huang, Cheng-Hong Wang et al. · 0 citations
Preprint Aug 2026

RVANNS: Mixed-Precision Indexing and Locality-Aware Graph Traversal on RISC-V

RVANNS is presented, an RVV-oriented ANNS engine that jointly optimizes vector representation and graph locality and achieves 3.39x and 4.94x speedups over scalar execution on real 128-bit and 256-bit RVV processors, respectively.

Cheng-Ying Huan, Yudong Liu, Jian-Guo Wang et al. · 0 citations
Jul 2026

Accelerating String-Heavy Queries with LLM Token Tables

Strings are the most common data type in modern database systems, yet they are often treated as an afterthought in high-performance data formats. While numerical data benefits from specialized, lightweight compression schemes, text is typically handled by general-purpose algorithms such as Zstd, LZ4, or Snappy, which r...

Tobias Schmidt, Nicolas Schmitt, Thomas Neumann et al. · 0 citations
Preprint Oct 2026

LLBPE: Linked-List Based GPU-Parallel BPE Tokenizer

Every LLM inference begins with tokenization, which converts raw input bytes into the discrete token sequence the model consumes. For text, this step is often implemented using Byte Pair Encoding (BPE), an algorithm originally introduced for data compression. BPE has traditionally run on the CPU with extensive optimiza...

Aditya Kovilur, V. C. Shekar, Ariful Azad · 0 citations
#machine learning Preprint Sep 2026

Hybrid GPU-CPU Retrieval for Personalized Search at Ultra-Large Scale

Embedding-based retrieval on user-generated content at the trillion-document scale exposes a sharp conflict between two production demands: deep, expressive personalization for queries with rich user intent, and broad coverage of a massive inventory under fixed latency and resource budgets. We characterize this as the...

Hao-Hao Fu, Ji-Chao Sun, Bai-Ting Zhu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.