Preprint
Aug 2026
CacheRoute: Planned Prefix-Affinity Routing for Large-Scale LLM Serving
When affinity recovers too little KV work, its residual load skew reduces or erases the improvement, so gating any deployment with a shadow replay rather than enabling affinity from workload statistics alone is recommended.
Huang Cheng
· 0 citations