Skip to content
Book Open access

Accelerating Graph-Based RAG Retrieval via Locality-Aware Device-Cloud Collaboration

Aug 2026 · Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 · pp. 839-850 · 0 citations · 14 references

Abstract

Retrieval-Augmented Generation (RAG) grounds large language models in external knowledge and has become a key technique for knowledge-intensive tasks. As knowledge bases continue to scale, however, the retrieval stage increasingly dominates end-to-end latency, limiting the responsiveness of RAG systems. In this paper, we identify and empirically validate a previously underexplored property of RAG workloads: strong per-user query locality, where individual users' queries concentrate on a small subset of the knowledge space. Motivated by this observation, we propose Lever, a locality-aware collaborative retrieval framework that exploits query locality to accelerate graph-based RAG retrieval. Lever maintains compact, personalized subgraph indexes on user's local devices as auxiliary structures to guide retrieval toward semantically relevant regions of a global index, enabling more efficient graph traversal without sacrificing coverage. To sustain effectiveness over time, Lever further incorporates adaptive resampling mechanisms that align on-device indexes with evolving query patterns. Extensive experiments on multiple RAG benchmarks demonstrate that Lever significantly reduces retrieval latency and improves throughput while preserving retrieval quality, highlighting query locality as a powerful and complementary lever for scalable RAG retrieval.

Read PDF

Similar papers

Book Open access Aug 2026

Structure Is All You Need to Reuse: Accelerating GraphRAG via Meta-Structure-Aware KV Caching

MetaKV is proposed, the first structure-aware KV caching mechanism that explicitly decouples static structural logic from dynamic entity semantics in GraphRAG inference, enabling high-throughput, low-latency GraphRAG without sacrificing adherence to graph topology.

Rui-Kun Luo, C. Gu, Jing Yang et al. · 0 citations
Open access 2026

Semantic Skyline: Multi-Embedding Skyline Query Processing for RAG Workloads in Vector Databases

Retrieval-Augmented Generation (RAG) systems and modern vector databases rely on embedding representations to retrieve semantically relevant entities. However, many real-world applications require multi-criteria query processing, where users express preferences across multiple semantic dimensions rather than a single n...

Md Arif Rahman, Taeyeon Kim, Shayhan Ameen Chowdhury et al. · 0 citations
Jul 2026

CGIF: Combining Proximity Graphs and Inverted Files for Efficient Filtered Vector Search over Arbitrary Predicates

Modern retrieval systems increasingly require filtered vector search under arbitrary predicate constraints, where users filter results by attributes such as category, price, location, keywords, and their combinations. Existing solutions either specialize in a single predicate type (e.g., range or equality filters), rel...

Jiarui Luo, Chaoji Zuo, Dong-Jie Deng · 0 citations
Conference Aug 2026

Query-Aware Multi-Relational Evidence Graph for Retrieval-Augmented Generation

Retrieval-Augmented Generation (RAG) systems typically apply identical retrieval logic to all user queries, overlooking that different query intents depend on fundamentally different document relationships. We propose QMEG, a query-aware retrieval framework that constructs a multi-relational evidence graph with five ac...

Jia-Run Pan, Yu-Ling Fan, Li Ma et al. · 0 citations
#artificial intelligence Preprint Sep 2026

MOSAIC: Query-Aware Exploration Policy Adaptation for GraphRAG

Graph Retrieval-Augmented Generation (GraphRAG) can connect evidence distributed across a corpus graph, but most systems use largely shared exploration procedures across queries. This creates a structural mismatch: direct facts may need compact local neighborhoods, comparisons need balanced coverage of multiple targets...

Eunkyeong Lee, Kyeong-Jin Oh, Jinwon Kim et al. · 0 citations
#artificial intelligence Preprint Sep 2026

LiteRAG: Cost-Efficient Graph-Based Retrieval-Augmented Generation

Graph-based retrieval can improve multi-hop question answering, but existing approaches often incur high query-time costs and produce diffuse, oversized contexts that reduce generation efficiency. We present LiteRAG, a graph-based retrieval method that replaces expensive retrieval-time LLM control with query-conditione...

Daniel Alejandro Coll Tejeda, Pedro García López, Daniel Barcelona-Pons · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.