Token Latency Fairness: Performance Isolation for Multi-Tenant LLM Serving
FairInference provides the novelelta-token fairness guarantee: for a well-behaved client, if a token is generated in d time units in isolation, it will be generated within d + {\delta} time units in multi-tenant execution, providing strong latency isolation guarantees for LLM serving.