CrossServe: Cross-Layer Scheduling for SLO Optimization in Multi-Tenant LLM Serving
The deployment of Large Language Models (LLMs) as multi-tenant cloud services is now widespread, but maintaining high Service Level Objective (SLO) attainment across diverse tenants remains challenging. Current serving systems focus on a single layer of the stack, either using iteration-level batching or coarse-grained...