Preprint
Aug 2026
An Internet for the KV Cache: Rethinking Classical Infrastructure Boundaries in the LLM Inference Age
This work argues that future inference infrastructure should allow decoupling of compute and KV Cache storage across cloud and datacenters, and proposes a vision for an Internet for the KV Cache, with KV Cache management working as a content-distribution system.
Siddhant Ray, Nick Feamster, Junchen Jiang
· 0 citations