Skip to content
Book Open access

Congestion-Aware Serving of Agentic LLM Applications

Sep 2026 · Proceedings of the 17th ACM SIGOPS Asia-Pacific Workshop on Systems · pp. 38-45 · 0 citations · 41 references

TL;DR

CALM-MAS is proposed, a congestion-aware serving framework for LLM applications that treats LLM test-time computation as an elastic resource, dynamically adjusting the compute profile of admitted tasks to tame congestion.

Abstract

Agentic LLM workflows issue many dependent calls with unpredictable resource demand, causing queue buildup and latency degradation on shared serving backends when left unmanaged. In this paper, we propose CALM-MAS, a congestion-aware serving framework for LLM applications that treats LLM test-time computation as an elastic resource, dynamically adjusting the compute profile of admitted tasks to tame congestion. CALM-MAS detects early signals of back-end saturation, and leverages the flexibility of LLM applications to regulate load. During spikes of requests, the system downgrades agent topology and reasoning depth; during low-utilization periods, it allocates additional reasoning effort to maximize task accuracy. Compared with a static serving baseline based on vLLM, CALM-MAS reduces shared-backend tail latency by 77% with the accuracy degradation remaining confined to 6.1 pps relative to the native agent configuration.

Read PDF

Similar papers

Preprint Sep 2026

PackServe: SLO-Aware Request Scheduling for Agentic LLM Serving at Scale

PackServe is a scheduler designed to reduce resource costs while meeting latency SLOs for agentic LLM serving, which uses compact white-box models to predict latency under prefill/decode interference and packs requests onto fewer serving instances while preserving KVC reuse and SLO constraints.

Zhi-Yuan Tan, De-Jiang Zhu, Jing-Zhe Jiang et al. · 0 citations
Preprint Sep 2026

SARA: SLO-Aware Resource Allocation for Disaggregated Agentic LLM Services

Recent advances in large language models (LLMs) are driving the emergence of multi-modal and agentic services for mobile users through cloud and edge infrastructures, where long-context workloads pose daunting challenges for inference latency. Existing disaggregated LLM serving systems largely rely on hardware profilin...

Shi-Cong Liu, Xiang-Hao Yu, Zheng-Run Gao et al. · 0 citations
Preprint Aug 2026

Bridging Agent Semantics with Spot Capacity: An Elastic and Recoverable Service Model

LLM agents increasingly drive long-running cloud inference workloads in which model calls differ in urgency, redundancy, completion semantics, and replay cost. Model-as-a-Service (MaaS) platforms expose several service models for trading cost against latency, availability, and capacity commitment. These models operate...

Min-Chen Yu · 0 citations
Preprint Aug 2026

From LLM Inference to Agentic Workloads: Characterization and Implications for Serving Systems

This work presents AgentSysBench, a benchmark suite and measurement toolkit with ten representative agentic applications and unified systems-level instrumentation, and identifies six properties that distinguish agentic workloads from conventional LLM serving.

Chaokun Chang, Yu-Kun Zhou, Kai-Hua Fu et al. · 10 citations
Open access Aug 2026

CELLServe: An SLO-Aware and Cost Efficient LLMs Serving System for Serverless Computing Environments

CELLServe formalizes SLO-constrained joint resource provisioning as an optimization problem with a dedicated algorithm, and introduces an opportunistic instance merging strategy for decode phase functions to reclaim fragmented resources.

Ze-Jian Wang, Nan Lin, Zi-Nuo Cai et al. · 0 citations
#artificial intelligence Preprint Sep 2026

End-to-End Latency-Minimizing and Load-Balanced Request Scheduling for Edge LLM Inference in Agentic AI Services

Large language model (LLM)-powered agentic AI services increasingly demand low-latency inference, motivating the deployment of LLMs across distributed edge servers. However, heterogeneous communication and computing capabilities, together with dynamically evolving inference states, make the edge server selection for ea...

Zhen Li, Jun Cai, Hao-Ran Gao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.