FairInference provides the novelelta-token fairness guarantee: for a well-behaved client, if a token is generated in d time units in isolation, it will be generated within d + {\delta} time units in multi-tenant execution, providing strong latency isolation guarantees for LLM serving.
Abstract
LLM serving is typically offered as a shared, multi-tenant service, where high-demand workloads from one client can cause latency SLO violations for others. Existing solutions for performance isolation equalize client throughput in the long run, for example through queueing and batching fairness. However, these approaches do not provide latency isolation guarantees; as a result, well-behaved clients can still experience significant degradation to their token-level latencies. In this paper, we present FairInference, which provides the novel {\delta}-token fairness guarantee: for a well-behaved client, if a token is generated in d time units in isolation, it will be generated within d + {\delta} time units in multi-tenant execution, providing strong latency isolation guarantees for LLM serving. To achieve this, FairInference addresses a key challenge of LLM serving: bounding delays from sharing GPU resources without support for fine-grained scheduling or resource allocation. In FairInference, the scheduler enforces per-token deadlines, while bounding the delays from GPU compute sharing and accounting for the additional delays introduced by the shared KV caching in GPU memory. We show that FairInference effectively bounds token-level latency spikes for well-behaved clients and improves overall throughput compared to state-of-the-art LLM serving systems.
RDPart is proposed, an OS-level Reuse-Driven LLC Partitioning policy designed to improve fairness while preserving the QoS of cloud workloads, and adopts a black-box design, making it well-suited for public cloud environments where real-time QoS feedback from applications is unavailable.
Javier Aznal, J. C. Saez, Carlos Bilbao· Proceedings of the Internati...· 1 citation
Large language models (LLMs) increasingly rely on context caching to enhance serving efficiency. However, this optimization inadvertently compromises fairness in multi-tenant LLM serving systems. Existing fair schedulers, which account only for compute resources, are unable to handle the multi-dimensional resource dema...
Zhuo-Yan Bai, Bin Gao, Fei Xu et al.· IEEE Transactions on Paralle...· 0 citations
The deployment of Large Language Models (LLMs) as multi-tenant cloud services is now widespread, but maintaining high Service Level Objective (SLO) attainment across diverse tenants remains challenging. Current serving systems focus on a single layer of the stack, either using iteration-level batching or coarse-grained...
Jia-He Li, Jia-Bin Li· Proceedings of the Internati...· 0 citations
A mathematical scheduling model that connects within-batch resource fairness to system throughput and provides a bi-criterion scheduling policy, ISJL, which maintains high throughput while aligning max-driven batch cost with token-metered revenue.
In this paper, we present a latency-aware scheduler for large-language-model (LLM) inference across mobile devices, edge servers, and a remote cloud. Our fine-grained delay model captures OFDMA uplink/downlink rates, KV-cache backhaul serialization, and profiled GPU planning-chunk resource constraints, enabling per-req...
Xinghan Wang, Xiao-Xiong Zhong, Wei-Hong Yang et al.· IEEE Transactions on Paralle...· 0 citations
Request scheduling is a key challenge in large-scale clusters serving agentic large language model (LLM) workloads. An effective scheduler must preserve key-value cache (KVC) reuse across long, shared prefixes, meet token-level latency service-level objectives (SLOs), and minimize GPU resource footprint. Existing sched...
Zhi-Yuan Tan, De-Jiang Zhu, Jing-Zhe Jiang et al.· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduOct 2, 2026
Martin Trust Center Managing Director Bill Aulet introduces Dear Dreamer, a free platform for middle and high school students who want to learn about entrepreneurship.
Microsoft Research Blog· microsoft.comSep 30, 2026
Extreme space-weather events can damage power systems on Earth and degrade GPS accuracy and satellite operations. A new machine learning system can predict where damage is likely to occur 30-60 minutes before a storm arrives. The post Forecasting space weather risks on power grids appeared first on Microsoft Research.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.