Skip to content
Book Open access

NumCache: KV Cache Compression and Retrieval for Financial Document QA

Aug 2026 · Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 · pp. 3655-3666 · 0 citations · 32 references

TL;DR

The proposed NumCache, which compresses SEC filings into KV caches initialized from numerically dense regions and trained directly on financial QAs, is evaluated, which highlights cache-based retrieval with number-preserving representations as an effective approach for long-context financial QA.

Abstract

Large Language Models (LLMs) are increasingly deployed in financial applications, particularly for interpreting U.S. Securities and Exchange Commission (SEC) filings. However, financial QA over these filings is challenging, as they are extremely long, numerically dense, and often require cross-document reasoning. Existing approaches struggle to scale to such settings due to long-context degradation and loss of numerical fidelity under context compression. Long-document Retrieval-Augmented Generation (RAG) improves evidence coverage through coarse-to-fine retrieval, yet semantic retrievers collapse fine-grained magnitudes and units, returning passages that lack the precise values needed for correct reasoning. Cache-Augmented Generation (CAG) projects attention states into compact KV representations and treats precomputed caches as reusable internal memory, but typically assumes caches are already well-formed and query-relevant, leaving open how to build and select number-faithful caches. To address this gap, we propose NumCache, which compresses SEC filings into KV caches initialized from numerically dense regions and trained directly on financial QAs. On top of these caches, we then train a contrastive retriever that aligns questions with cache representations, thus improving retrieval performance. We evaluate NumCache on the Fin-RATE benchmark, which comprises SEC-filing QA tasks covering single-filing reasoning, cross-firm comparison, and longitudinal trend analysis. NumCache achieves up to 4× context compression with competitive accuracy (41.4% vs.\ 44.8% for uncompressed Qwen3-4B on single-filing reasoning) and a 6.4× inference speedup, while its contrastive retriever attains 68.3% Recall@1 on single-filing retrieval; 2.3× the strongest text-based baseline (29.7%). These results highlight cache-based retrieval with number-preserving representations as an effective approach for long-context financial QA.

Read PDF

Similar papers

#machine learning Preprint Sep 2026

Cartridges++: KV Cache Compression without Off-Context Derailment

Serving long documents to a Large Language Model (LLM) repeatedly is expensive: computations grow with context length, and the memory footprint of the key-value (KV) cache balloons. Compressed KV (CKV) representations aim to mimic the cache of a document and are typically computed once and for all, ahead of inference t...

Sonia Laguna, João Monteiro, Marco Cuturi et al. · 0 citations
#machine learning Preprint Sep 2026

REVA: Reusable Evidence View Aggregation for Context-Efficient RAG Serving

Retrieval-augmented generation (RAG) improves knowledge-intensive large language model (LLM) applications by conditioning generation on retrieved documents, but longer contexts increase latency, key-value (KV) cache memory, and token cost. Post-retrieval compression can reduce this cost, yet existing compressors often...

T. Nguyen, Qi-Ran Hu, Ban-Ruo Liu et al. · 0 citations
#machine learning Preprint Sep 2026

PatchKV: Weight-Space Compensation of KV Cache

Long-context inference with Large Language Models (LLMs) is bottlenecked by the linearly growing memory of the key-value (KV) cache. Existing compression methods reduce the cache through token eviction or approximation, but degrade sharply at aggressive compression budgets. We propose PatchKV, a training-free framework...

Chanryeol Lee, Chanhyuk Lee, Yeonwoo Choi et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Where Should a Document Live: Context, Representations, or Parameters?

To answer questions outside of their pre-training data, large language models (LLMs) need access to new information, which can be presented in the context window as documents, encoded into the model's parameters, or injected as latent representations. However, each of these methods comes with different efficiency, cost...

Nathanaël Carraz Rakotonirina, Momchil Hardalov, Gonzalo Iglesias et al. · 0 citations
Preprint Aug 2026

PuzzleKV: Page-Wise Low-Rank Decomposition for KV Cache Compression

Long-context inference in large language models (LLMs) is increasingly limited by the memory required for the key-value (KV) cache. KV cache compression addresses this problem by reducing the storage cost of previous tokens. Among existing approaches, low-rank compression is particularly attractive because it represent...

Zi-Zhong Wang, Jie-Ying Wang, Zhao Zhang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Fine-Tuning a KV Cache Concatenation-Aware Model or Recomputing KV Caches? Why Not Both?

In Retrieval-Augmented Generation (RAG) systems, a large number of retrieved chunks are concatenated to form the input context so that users can receive high-quality responses based on external knowledge. As a result, the input context length increases substantially, leading to a larger prefill workload and, in turn, a...

F. Tachibana, Daisuke Miyashita, Jun Deguchi · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.