Sep 2026· Proceedings of the 18th ACM Workshop on Hot Topics in Storage and File Systems· 0 citations· 11 references
TL;DR
The challenges of the restoration-recomputation trade-off are investigated and its impact on inference performance when left unaddressed, and an I/O-aware KV-cache management policy is presented that dynamically navigates this trade-off.
Abstract
Key-Value (KV)-caches are essential to modern Large Language Model (LLM) inference, transforming inference from a computationally prohibitive task into a practical one. However, the limited capacity of GPU High Bandwidth Memory (HBM) is often insufficient to retain KV-cache data for all active requests, especially when large models and long contexts are served. To address this limitation, modern inference systems support offloading KV-cache data from HBM to CPU DRAM, local, or remote storage. A known issue is the trade-off between restoring offloaded KV-cache data versus recomputing; the right policy is non-trivial and depends on dynamic factors such as storage bandwidth, interconnect performance, workload characteristics, context length, and service-level objectives (SLOs). In this paper, we first investigate the challenges of the restoration-recomputation trade-off and quantify its impact on inference performance when left unaddressed. We then present an I/O-aware KV-cache management policy that dynamically navigates this trade-off. Our approach maximizes inference throughput while satisfying service-level objectives by either restoring or recomputing KV-cache blocks based on the performance characteristics of the GPU and storage tiers. Initial theoretical results show a 10 × improvement in performance over simplistic static policies while maintaining SLO compliance.
This model reveals one key opportunity: dividing a restore request proportionally between the storage path and the GPU can improve inference performance while still meeting SLOs, and reduces the KV-cache storage stack to a performance model based on per-tier capacity, per-tier and interconnect bandwidth, and GPU arithm...
OasisKV is presented, a memory-centric LLM inference system design that alleviates HBM capacity pressure by decoupling full KV-cache storage from HBM during LLM decoding and observes that future important tokens can be predicted accurately in advance using lookahead tokens drafted by speculative decoding (SD).
Key--value (KV) cache compression is an effective way to reduce the memory overhead of large language model (LLM) inference, particularly for long-context workloads. However, existing compression methods make different trade-offs among accuracy, inference latency, and peak KV cache memory utilization, making a single f...
Michael Wang, Keith Li, Roozbeh Bostandoost· 0 citations
KV-cache offload is widely used to stretch GPU memory for LLM serving, but its storage behavior has not been characterized at the block-device level. In this paper, we study LMCache through realworld multi-session workloads that span same/different context $\times$ same/different prompt, using over 100 stateless reques...
Ying He, Dingsen Shi, Yanbo Dai et al.· 2026 International Conferenc...· 0 citations
Disaggregated LLM serving separates prefill and decode into distinct node pools, interposing a network fabric between the moment a key-value (KV) cache is computed and the moment it is consumed. This architectural shift invalidates a core assumption of classical cache policies: that the cost of a miss is simply recompu...
Dong Liu, Yan-Xuan Yu, Eric Jiang et al.· Proceedings of the 19th ACM...· 0 citations
It is found that cache performance depends on transfer granularity, intermediate memory use, and when transfers enter the request schedule, not only on device bandwidth, not only on device bandwidth.
Joseph Kanichai, T. De Matteis, Animesh Trivedi· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.