Aug 2026· Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2· pp. 839-850· 0 citations· 14 references
Abstract
Retrieval-Augmented Generation (RAG) grounds large language models in external knowledge and has become a key technique for knowledge-intensive tasks. As knowledge bases continue to scale, however, the retrieval stage increasingly dominates end-to-end latency, limiting the responsiveness of RAG systems. In this paper, we identify and empirically validate a previously underexplored property of RAG workloads: strong per-user query locality, where individual users' queries concentrate on a small subset of the knowledge space. Motivated by this observation, we propose Lever, a locality-aware collaborative retrieval framework that exploits query locality to accelerate graph-based RAG retrieval. Lever maintains compact, personalized subgraph indexes on user's local devices as auxiliary structures to guide retrieval toward semantically relevant regions of a global index, enabling more efficient graph traversal without sacrificing coverage. To sustain effectiveness over time, Lever further incorporates adaptive resampling mechanisms that align on-device indexes with evolving query patterns. Extensive experiments on multiple RAG benchmarks demonstrate that Lever significantly reduces retrieval latency and improves throughput while preserving retrieval quality, highlighting query locality as a powerful and complementary lever for scalable RAG retrieval.
MetaKV is proposed, the first structure-aware KV caching mechanism that explicitly decouples static structural logic from dynamic entity semantics in GraphRAG inference, enabling high-throughput, low-latency GraphRAG without sacrificing adherence to graph topology.
Rui-Kun Luo, C. Gu, Jing Yang et al.· Proceedings of the 32nd ACM...· 0 citations
Retrieval-Augmented Generation (RAG) systems and modern vector databases rely on embedding representations to retrieve semantically relevant entities. However, many real-world applications require multi-criteria query processing, where users express preferences across multiple semantic dimensions rather than a single n...
Md Arif Rahman, Taeyeon Kim, Shayhan Ameen Chowdhury et al.· IEEE Access· 0 citations
Modern retrieval systems increasingly require filtered vector search under arbitrary predicate constraints, where users filter results by attributes such as category, price, location, keywords, and their combinations. Existing solutions either specialize in a single predicate type (e.g., range or equality filters), rel...
Jiarui Luo, Chaoji Zuo, Dong-Jie Deng· Proceedings of the VLDB Endo...· 0 citations
Retrieval-Augmented Generation (RAG) systems typically apply identical retrieval logic to all user queries, overlooking that different query intents depend on fundamentally different document relationships. We propose QMEG, a query-aware retrieval framework that constructs a multi-relational evidence graph with five ac...
Jia-Run Pan, Yu-Ling Fan, Li Ma et al.· 2026 3rd International Confe...· 0 citations
Graph Retrieval-Augmented Generation (GraphRAG) can connect evidence distributed across a corpus graph, but most systems use largely shared exploration procedures across queries. This creates a structural mismatch: direct facts may need compact local neighborhoods, comparisons need balanced coverage of multiple targets...
Eunkyeong Lee, Kyeong-Jin Oh, Jinwon Kim et al.· 0 citations
Graph-based retrieval can improve multi-hop question answering, but existing approaches often incur high query-time costs and produce diffuse, oversized contexts that reduce generation efficiency. We present LiteRAG, a graph-based retrieval method that replaces expensive retrieval-time LLM control with query-conditione...
Daniel Alejandro Coll Tejeda, Pedro García López, Daniel Barcelona-Pons· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.