Preprint
Jul 2026
InferScale: GPU-Native KV Injection for Personalized LLM Serving
This work presents InferScale, a GPU-native LLM memory system that replaces repeated prompt prefilling with reusable KV state, and encodes each memory fact together with a small window of preceding conversation context while caching only the target fact's KV.
Peter Li, Prashant Pandey
· 1 citation