Adaptive GPU Sharing for Real-time LLM Serving with Best-effort Workloads
Large Language Models (LLMs) are increasingly deployed in latency-sensitive applications, where real-time serving must satisfy stringent service-level objectives (SLOs). However, request intensities fluctuate over time, and under low load LLM services leave a substantial fraction of GPU task idle. A promising approach...