Congestion-Aware Serving of Agentic LLM Applications
CALM-MAS is proposed, a congestion-aware serving framework for LLM applications that treats LLM test-time computation as an elastic resource, dynamically adjusting the compute profile of admitted tasks to tame congestion.