This paper proposes a vision for telemetry-informed adaptive compression in edge RAG, grounded in experimental evidence, and argues for runtime policies that dynamically manage compression, guided by workload features and edge telemetry.
Abstract
Retrieval-augmented generation (RAG) improves language-model responses by grounding generation in external passages, which comes with overhead: retrieved context lengthens the prompt, increasing prefill work, KV-cache footprint, memory traffic, latency, and energy. Context compression offers a natural remedy by pruning retrieved text before generation. However, state-of-the-art context-compression methods are typically used with a fixed compression budget, or with the rate selected offline and then applied at inference time. This static view ignores both workload variation and the live state of the edge device. On an edge SoC, compression is not free: the compressor itself runs on the same SoC and consumes latency and energy that can offset any generation savings. This paper proposes a vision for telemetry-informed adaptive compression in edge RAG, grounded in experimental evidence. We characterize the compression tradeoff on the NVIDIA Jetson AGX Thor using Llama and Qwen generators, Natural Questions and HotpotQA datasets, and LLMLingua-2 compression. Our measurements show that generation dominates the RAG budget for larger models, reaching roughly 90% of per-query latency and 91% of GPU energy for 7B-8B generators. Exploring the impact of the compression rate reveals an adaptive operating region: mild compression can miss energy opportunities, and overly aggressive compression can hurt inference quality. Intermediate compression can reduce GPU energy by up to 53.2%, and SoC energy by up to 48.2%, with negligible quality loss. We argue for runtime policies that dynamically manage compression, guided by workload features and edge telemetry.
Serving long documents to a Large Language Model (LLM) repeatedly is expensive: computations grow with context length, and the memory footprint of the key-value (KV) cache balloons. Compressed KV (CKV) representations aim to mimic the cache of a document and are typically computed once and for all, ahead of inference t...
Sonia Laguna, João Monteiro, Marco Cuturi et al.· 1 citation
This work revisits RAG compression from a data-mining perspective by aggregating historical query--document--model interactions into reusable evidence views and proposes Reusable Evidence View Aggregation (REVA), a framework that mines the target generator's historical attention traces into a document-keyed, budget-agn...
T. Nguyen, Qi-Ran Hu, Ban-Ruo Liu et al.· 0 citations
This work learns a selector that is able to produce, based on context, a per-layer cache configuration towards an overall compression budget, and achieves the Pareto frontier of accuracy against cache size, against learning-free and post-hoc baselines.
João Monteiro, Louis Béthune, A. Filippova et al.· 1 citation
Retrieval-Augmented Generation (RAG) has become a standard paradigm for grounding large language models (LLMs) in external knowledge, but typical deployments introduce substantial energy, latency, and cost overhead due to expensive retrieval and context-processing pipelines. Recent work in "Green AI" and sustainable ma...
Anupam Dhakal, Prashant Pokharel, S. Adhikari· European Journal of Applied...· 0 citations
We present RAGMark, a modular benchmarking framework for advanced Retrieval-Augmented Generation (RAG) systems targeting small-scale multi-GPU environments. RAGMark evaluates diverse RAG components, including retrievers, vector databases, prompt-processing methods, and generator models, while collecting detailed per-st...
Z. Feric, Amir Taherin, Bin Ren et al.· 0 citations
KV caches are a major bottleneck in long-context inference and long-form generation with large language models. Existing training-free eviction policies largely rely on proxy importance signals, such as attention mass, to decide which past tokens to retain. We argue that cache compression should instead preserve the pr...
Doo Hwan Hwang, Junyoung Jang, Jun-Ho Na et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.