#machine learning
Feb 2026
PrefillShare: A Shared Prefill Module for KV Reuse in Multi-LLM Disaggregated Serving
This work proposes PrefillShare, a novel algorithm that enables sharing the prefill stage across multiple fine-tuned models in a disaggregated setting and introduces a routing mechanism that enables effective prefill sharing in a vLLM-based disaggregated system.
Sunghyeon Woo, Hoseung Kim, S. Shim et al.
· arXiv.org · 3 citations