Preprint
Jul 2026
MemDecay: Region-Aware KV Cache Eviction for Efficient LLM Agent Inference
These results establish semantic prompt structure as a robust signal for KV-cache management while clarifying how it should be combined with attention-based importance.
Venkatesha Matam, Keon Kim
· 1 citation