Skip to content

Token Latency Fairness: Performance Isolation for Multi-Tenant LLM Serving

Sep 2026 · 0 citations · 40 references
Computer Science

TL;DR

FairInference provides the novelelta-token fairness guarantee: for a well-behaved client, if a token is generated in d time units in isolation, it will be generated within d + {\delta} time units in multi-tenant execution, providing strong latency isolation guarantees for LLM serving.

Abstract

LLM serving is typically offered as a shared, multi-tenant service, where high-demand workloads from one client can cause latency SLO violations for others. Existing solutions for performance isolation equalize client throughput in the long run, for example through queueing and batching fairness. However, these approaches do not provide latency isolation guarantees; as a result, well-behaved clients can still experience significant degradation to their token-level latencies. In this paper, we present FairInference, which provides the novel {\delta}-token fairness guarantee: for a well-behaved client, if a token is generated in d time units in isolation, it will be generated within d + {\delta} time units in multi-tenant execution, providing strong latency isolation guarantees for LLM serving. To achieve this, FairInference addresses a key challenge of LLM serving: bounding delays from sharing GPU resources without support for fine-grained scheduling or resource allocation. In FairInference, the scheduler enforces per-token deadlines, while bounding the delays from GPU compute sharing and accounting for the additional delays introduced by the shared KV caching in GPU memory. We show that FairInference effectively bounds token-level latency spikes for well-behaved clients and improves overall throughput compared to state-of-the-art LLM serving systems.

View source

Similar papers

Book Open access Sep 2026

RDPart: A reuse-based OS-level cache-partitioning policy for fairness optimization in cloud data centers

RDPart is proposed, an OS-level Reuse-Driven LLC Partitioning policy designed to improve fairness while preserving the QoS of cloud workloads, and adopts a black-box design, making it well-suited for public cloud environments where real-time QoS feedback from applications is unavailable.

Javier Aznal, J. C. Saez, Carlos Bilbao · 1 citation

FairCache: Demystifying Cache-Induced Unfairness in Multi-Tenant Large Language Model Serving

Large language models (LLMs) increasingly rely on context caching to enhance serving efficiency. However, this optimization inadvertently compromises fairness in multi-tenant LLM serving systems. Existing fair schedulers, which account only for compute resources, are unable to handle the multi-dimensional resource dema...

Zhuo-Yan Bai, Bin Gao, Fei Xu et al. · 0 citations
Book Open access Sep 2026

CrossServe: Cross-Layer Scheduling for SLO Optimization in Multi-Tenant LLM Serving

The deployment of Large Language Models (LLMs) as multi-tenant cloud services is now widespread, but maintaining high Service Level Objective (SLO) attainment across diverse tenants remains challenging. Current serving systems focus on a single layer of the stack, either using iteration-level batching or coarse-grained...

Jia-He Li, Jia-Bin Li · 0 citations
Preprint Aug 2026

Efficiency and Cost Alignment in Batched LLM Serving via Resource-Fair Scheduling

A mathematical scheduling model that connects within-batch resource fairness to system throughput and provides a bi-criterion scheduling policy, ISJL, which maintains high throughput while aligning max-driven batch cost with token-metered revenue.

Da-Yi Yao, Zijie Zhou · 0 citations

Mix-or-Split: Latency-Aware Scheduling for Edge–Cloud LLM Inference

In this paper, we present a latency-aware scheduler for large-language-model (LLM) inference across mobile devices, edge servers, and a remote cloud. Our fine-grained delay model captures OFDMA uplink/downlink rates, KV-cache backhaul serialization, and profiled GPU planning-chunk resource constraints, enabling per-req...

Xinghan Wang, Xiao-Xiong Zhong, Wei-Hong Yang et al. · 0 citations
Preprint Sep 2026

PackServe: SLO-Aware Request Scheduling for Agentic LLM Serving at Scale

Request scheduling is a key challenge in large-scale clusters serving agentic large language model (LLM) workloads. An effective scheduler must preserve key-value cache (KVC) reuse across long, shared prefixes, meet token-level latency service-level objectives (SLOs), and minimize GPU resource footprint. Existing sched...

Zhi-Yuan Tan, De-Jiang Zhu, Jing-Zhe Jiang et al. · 0 citations

Related blog posts

Microsoft Research Blog Sep 30, 2026

Forecasting space weather risks on power grids

Extreme space-weather events can damage power systems on Earth and degrade GPS accuracy and satellite operations. A new machine learning system can predict where damage is likely to occur 30-60 minutes before a storm arrives. The post Forecasting space weather risks on power grids appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.