Skip to content
Book Open access

NumCache: KV Cache Compression and Retrieval for Financial Document QA

Aug 2026 · Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 · pp. 3655-3666 · 0 citations · 11 references

Abstract

Large Language Models (LLMs) are increasingly deployed in financial applications, particularly for interpreting U.S. Securities and Exchange Commission (SEC) filings. However, financial QA over these filings is challenging, as they are extremely long, numerically dense, and often require cross-document reasoning. Existing approaches struggle to scale to such settings due to long-context degradation and loss of numerical fidelity under context compression. Long-document Retrieval-Augmented Generation (RAG) improves evidence coverage through coarse-to-fine retrieval, yet semantic retrievers collapse fine-grained magnitudes and units, returning passages that lack the precise values needed for correct reasoning. Cache-Augmented Generation (CAG) projects attention states into compact KV representations and treats precomputed caches as reusable internal memory, but typically assumes caches are already well-formed and query-relevant, leaving open how to build and select number-faithful caches. To address this gap, we propose NumCache, which compresses SEC filings into KV caches initialized from numerically dense regions and trained directly on financial QAs. On top of these caches, we then train a contrastive retriever that aligns questions with cache representations, thus improving retrieval performance. We evaluate NumCache on the Fin-RATE benchmark, which comprises SEC-filing QA tasks covering single-filing reasoning, cross-firm comparison, and longitudinal trend analysis. NumCache achieves up to 4× context compression with competitive accuracy (41.4% vs.\ 44.8% for uncompressed Qwen3-4B on single-filing reasoning) and a 6.4× inference speedup, while its contrastive retriever attains 68.3% Recall@1 on single-filing retrieval; 2.3× the strongest text-based baseline (29.7%). These results highlight cache-based retrieval with number-preserving representations as an effective approach for long-context financial QA.

Read PDF

Similar papers

#machine learning Preprint Sep 2026

REVA: Reusable Evidence View Aggregation for Context-Efficient RAG Serving

Retrieval-augmented generation (RAG) improves knowledge-intensive large language model (LLM) applications by conditioning generation on retrieved documents, but longer contexts increase latency, key-value (KV) cache memory, and token cost. Post-retrieval compression can reduce this cost, yet existing compressors often...

T. Nguyen, Qi-Ran Hu, Ban-Ruo Liu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Where Should a Document Live: Context, Representations, or Parameters?

To answer questions outside of their pre-training data, large language models (LLMs) need access to new information, which can be presented in the context window as documents, encoded into the model's parameters, or injected as latent representations. However, each of these methods comes with different efficiency, cost...

Nathanaël Carraz Rakotonirina, Momchil Hardalov, Gonzalo Iglesias et al. · 0 citations
Preprint Aug 2026

PuzzleKV: Page-Wise Low-Rank Decomposition for KV Cache Compression

Long-context inference in large language models (LLMs) is increasingly limited by the memory required for the key-value (KV) cache. KV cache compression addresses this problem by reducing the storage cost of previous tokens. Among existing approaches, low-rank compression is particularly attractive because it represent...

Zi-Zhong Wang, Jie-Ying Wang, Zhao Zhang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Fine-Tuning a KV Cache Concatenation-Aware Model or Recomputing KV Caches? Why Not Both?

In Retrieval-Augmented Generation (RAG) systems, a large number of retrieved chunks are concatenated to form the input context so that users can receive high-quality responses based on external knowledge. As a result, the input context length increases substantially, leading to a larger prefill workload and, in turn, a...

F. Tachibana, Daisuke Miyashita, Jun Deguchi · 0 citations
Preprint Aug 2026

MegaMem: A Retrieval Solution for Ultra-Large Context Windows

These results show that MegaMem supports ultra-large persistent memory while preserving strong answer accuracy under a bounded generation context, and provides a practical path toward accurate retrieval over memories ranging from hundreds of millions to one billion tokens.

Xin-Yuan Song, Bo-Wen Zhu, H. Haque et al. · 0 citations
Jul 2026

SemPIC: Learning Semantic Position-Independent KV Caches

This work presents SemPIC, which trains a LoRA-enabled Writer to compile native per-layer document KVs through behavioral distillation while retaining the pretrained decoder as an unchanged Reader, and introduces KV Gradient Checkpointing, which reduces peak training memory without severing gradients through cached KVs...

Hui Xie, Peng Xiao, Yutong Deng et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.