AsymFlow: Enabling Long-Context LLM Serving via CPU-GPU Prefill-Decode Disaggregation
Driven by cost and privacy constraints, many organizations deploy large language model (LLM) inference on local CPU-GPU servers. For long-context requests, the key-value (KV) cache grows linearly with sequence length and can rapidly exhaust GPU memory, reducing both the maximum supported context and request concurrency...