This work shows that the operations that such knowledge bases allow can be replicated with zero ingestion costs (not even a vector database), and introduces Limited-Ingestion ScalableRAG, which does use a minimal vector database as well as an automated pattern discovery from a sample of documents, to further improve accuracy at scale.
Abstract
Recent advances in RAG aim to optimize for performance by paying high ingestion costs for knowledge ingestion: building knowledge graphs or extracting SQL tables. In this work we show that the operations that such knowledge bases allow can be replicated with zero ingestion costs (not even a vector database); in fact our solution, Zero-Ingestion ScalableRAG, handily out-performs all baselines (including knowledge graph approaches) in three out of the six corpora considered here, and only marginally missing maximum performance on the other three, with average accuracy across all six datasets 7.36% above the next most competitive baseline. It achieves this by keeping a workspace of document sets and values sets that it can write into and read from, allowing for on-the-fly aggregative reasoning in all situations where grouping is required on a primary key that is in one to one correspondence with a subset of the total document set. Capping the number of LLM calls by a constant independent of the corpus size, we also introduce Limited-Ingestion ScalableRAG, which does use a minimal vector database as well as an automated pattern discovery from a sample of documents, to further improve accuracy at scale. Our code is available at https://github.com/cohesity/ScalableRAG .
This work compares the trade-offs between retrieval-augmented generation, prefix caching, and fine-tuning in large language models to show that when the knowledge is stable and matches the questions asked, fine-tuning is the greenest, fastest, and most accurate, because it produces shorter answers.
Serving long documents to a Large Language Model (LLM) repeatedly is expensive: computations grow with context length, and the memory footprint of the key-value (KV) cache balloons. Compressed KV (CKV) representations aim to mimic the cache of a document and are typically computed once and for all, ahead of inference t...
Sonia Laguna, João Monteiro, Marco Cuturi et al.· 1 citation
Graph with Adaptive Shortcuts (GAS), a framework that leverages historical query logs to build lightweight auxiliary structures, enhancing search efficiency over a single base graph with minimal overhead, and consistently outperforms existing general indexes in wide-table scenarios.
Zi-Yuan He, Yu-Xiang Wang, Yu Sun et al.· Proceedings of the VLDB Endo...· 0 citations
Generating executable structured queries for OpenStreetMap (OSM) from natural language remains a challenging task due to the rigid syntax of OverpassQL and the steep learning curve it presents to end-users. While steering Large Language Models (LLMs) via
demonstration contexts
offers a promising solution, existing...
Zhuoyue Wan, Wentao Hu, Hwanhee Kim et al.· Proceedings of the ACM on Ma...· 0 citations
A single model scale challenges the flexibility of a production retrieval system: some settings need it faster, others need a smaller index, and the right trade-off changes with the workload. In the context of information retrieval (IR), a transformer-based model can be made smaller in three ways---using fewer layers,...
Yu Wang, Shengyao Zhuang, Xue-Guang Ma et al.· 1 citation
SQL is the database community's success story in terms of language design. The key reason for its success is its declarativeness: it gives rise to optimizability, reducing the programmer's burden significantly. However, given the evolving complexity of problems to solve with query languages, our community needs to re-t...
Molham Aref, Leonid Libkin, Wim Martens· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.