Large language model (LLM) agents execute applications through a workflow of inference requests with tool calls and user interactions. Serving these applications at production scale requires understanding how application behavior shapes inference demand and for guiding efficient execution. Recent characterization studi...
Yi-Hao Zheng, Jing-Zhe Jiang, De-Jiang Zhu et al.· 0 citations
Request scheduling is a key challenge in large-scale clusters serving agentic large language model (LLM) workloads. An effective scheduler must preserve key-value cache (KVC) reuse across long, shared prefixes, meet token-level latency service-level objectives (SLOs), and minimize GPU resource footprint. Existing sched...
Zhi-Yuan Tan, De-Jiang Zhu, Jing-Zhe Jiang et al.· 0 citations
Online image generation with Diffusion Transformers (DiTs) must meet latency service-level objectives (SLOs) while using GPU resources efficiently. Existing systems improve GPU utilization by batching multiple requests for joint execution. However, request-level batching offers limited control over batch size: batches...
Zhexiang Zhang, Min-Chen Yu, Yi-Fan Sun et al.· 0 citations
LLM agents increasingly drive long-running cloud inference workloads in which model calls differ in urgency, redundancy, completion semantics, and replay cost. Model-as-a-Service (MaaS) platforms expose several service models for trading cost against latency, availability, and capacity commitment. These models operate...
Min-Chen Yu· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.