Book
Open access
Aug 2026
DualPath: Accelerating Agentic LLM Inference by Harvesting Disaggregated KV-Cache Storage I/O
DualPath is an inference system that breaks this bottleneck by introducing dual-path KV-Cache loading and enables a novel storage-to-decode path, in which the KV-Cache is loaded into decoding engines and then efficiently transferred to prefill engines via RDMA over the compute network.
Yongtong Wu, Shaoyuan Chen, Rilin Huang et al.
· Proceedings of the ACM SIGCO... · 0 citations