Skip to content
Book Open access

Structure Is All You Need to Reuse: Accelerating GraphRAG via Meta-Structure-Aware KV Caching

Aug 2026 · Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 · 0 citations · 40 references

TL;DR

MetaKV is proposed, the first structure-aware KV caching mechanism that explicitly decouples static structural logic from dynamic entity semantics in GraphRAG inference, enabling high-throughput, low-latency GraphRAG without sacrificing adherence to graph topology.

Abstract

Retrieval-Augmented Generation over Knowledge Graphs (GraphRAG) enhances Large Language Models (LLMs) with structured, multi-hop evidence. However, existing GraphRAG systems predominantly linearize retrieved subgraphs into long textual prompts, forcing LLMs to recompute identical schema-level reasoning across queries repeatedly. This text-centric design incurs substantial prefilling latency, memory overhead, and severely limited cache reuse under entity-level variations. We observe that although retrieved entities differ across queries, their underlying logical schemas (meta-structures) recur with high frequency, indicating that most computational cost is spent on repeatedly encoding invariant structural logic. In this paper, we propose MetaKV, the first structure-aware KV caching mechanism that explicitly decouples static structural logic from dynamic entity semantics in GraphRAG inference. In a preparation phase, MetaKV mines frequent meta-structures and pre-computes their Key-Value (KV) caches as reusable Skeleton KVs. During inference, query-specific entity representations are injected into reserved structural slots to assemble the context without recomputing graph topology. To further enforce faithfulness to graph reasoning, MetaKV introduces a Topological Mask that constrains attention to valid graph edges. Extensive experiments conducted on HotpotQA and MetaQA datasets demonstrate that MetaKV achieves up to 6.4× prefilling speedup and a 73% effective cache-hit rate while maintaining competitive reasoning accuracy, enabling high-throughput, low-latency GraphRAG without sacrificing adherence to graph topology.

Read PDF

Similar papers

Preprint Aug 2026

KGCache: Amortized Subgraph Retrieval for KG Reasoning with LLMs

This work proposes KGCache, an in-memory cache for one-hop knowledge graph neighborhoods, which is designed to be compatible with both iterative traversal (ToG) and one shot planning (RoG) KGQA paradigms and shows substantial entity reuse among starting entities and entities reached during traversal.

Uros Stanic, Chang-He Yuan, Sabuj Laskar et al. · 0 citations
Book Open access Aug 2026

Accelerating Graph-Based RAG Retrieval via Locality-Aware Device-Cloud Collaboration

This paper proposes Lever, a locality-aware collaborative retrieval framework that exploits query locality to accelerate graph-based RAG retrieval, and identifies and empirically validate a previously underexplored property of RAG workloads: strong per-user query locality.

Yong-Heng Deng, Tianyuan Jiang, Zhen-Ya Ma et al. · 0 citations
#machine learning Preprint Sep 2026

REVA: Reusable Evidence View Aggregation for Context-Efficient RAG Serving

Retrieval-augmented generation (RAG) improves knowledge-intensive large language model (LLM) applications by conditioning generation on retrieved documents, but longer contexts increase latency, key-value (KV) cache memory, and token cost. Post-retrieval compression can reduce this cost, yet existing compressors often...

T. Nguyen, Qi-Ran Hu, Ban-Ruo Liu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

MOSAIC: Query-Aware Exploration Policy Adaptation for GraphRAG

Graph Retrieval-Augmented Generation (GraphRAG) can connect evidence distributed across a corpus graph, but most systems use largely shared exploration procedures across queries. This creates a structural mismatch: direct facts may need compact local neighborhoods, comparisons need balanced coverage of multiple targets...

Eunkyeong Lee, Kyeong-Jin Oh, Jinwon Kim et al. · 0 citations
#artificial intelligence Preprint Sep 2026

RelaxKV: Recomputation Guided by the Query with Sparse Context Attention for Efficient KV Cache Reuse

Cross-request KV caching reduces the prefill cost of Retrieval-Augmented Generation (RAG), but conventional prefix caching severely limits cache reuse across requests. Position-Independent Caching (PIC) removes this constraint by reusing independent chunks, but their KV states miss cross-chunk interactions. Existing me...

Ruo-Ling Qi, Yi-Rui Liu, Xuan'er Wu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.