An Approximate Queueing Model of LLM Inference Serving for SLO-Driven Autoscaling
Performance models of LLM servers support both latency evaluation and the design of controllers for autoscaling against service level objectives (SLOs) and for inference optimization. We model the multiplexed execution of prefill and decode operations with a tractable, approximate queueing model under Markovian assumpt...