Conference
Open access
Jul 2026
Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization
This work systematizes a rapidly evolving area, providing a foundation for understanding and innovating KV cache designs in modern LLM serving infrastructure.
Jiantong Jiang, Peiyu Yang, Rui Zhang et al.
· Annual Meeting of the Associ... · 16 citations
· ⚡1