Preprint
Jul 2026
A Year in LLM Serving: Workload Evolution, Caching and Load-Balancing
This work further the understanding of real-world LLM serving workloads through both a global characterization and a longitudinal study of a one-year production trace from Chutes, revealing workload evolution and user-model structure that are typically hidden behind aggregate views.
William Nixon, Jon Durbin, Florian Standhartinger et al.
· 0 citations