In this paper, we study a mixed-prompt scenario—where both short and long prompts coexist—in an LLM inference serving system that supports diverse applications with heterogeneous iteration-time SLOs. To improve throughput for long prompts, prior work divides them into chunks and batches requests or chunks to meet the t...
Hai-Ying Shen, Tanmoy Sen, Yuxiong He· Proceedings of the Internati...· 0 citations
Large Language Model (LLM) serving systems face a KV-cache (KVC) bottleneck. In this paper, our experimental study shows that block-based allocation increases Time-Between-Tokens (TBT) due to preemptions, while prediction-based allocation increases Time-to-First-Token (TTFT) and TBT due to allocated but unused KVC and...
Haiying Shen, Tanmoy Sen, Masahiro Tanaka· International Conference on...· 0 citations
Public LLM services serve diverse multi-agent applications with varying workflow dependencies and performance requirements. Requests generated by these applications often exhibit commonality and interdependence, yet current systems largely ignore such application-level structure. As a result, at the LLM engine cluster...
Uttam Rao, Ali Zafar Sadiq, Hai-Ying Shen et al.· International Conference on...· 0 citations
Multi-agent applications increasingly rely on shared large language model backends in the public cloud, where bursty workloads cause requests from different agents to contend for the same LLM instances, leading to long queues, memory imbalance, and severe tail-latency inflation. Existing approaches typically prioritize...
Ali Zafar Sadiq, Hai-Ying Shen· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.