Skip to content

Author

Tanmoy Sen

We have 2 of 37 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Book Open access Sep 2026

Heterogeneous SLO Guaranteed Multi-Resource-Aware Batching in LLM Serving

In this paper, we study a mixed-prompt scenario—where both short and long prompts coexist—in an LLM inference serving system that supports diverse applications with heterogeneous iteration-time SLOs. To improve throughput for long prompts, prior work divides them into chunks and batches requests or chunks to meet the t...

Hai-Ying Shen, Tanmoy Sen, Yuxiong He · 0 citations
Conference Jul 2026

Managing KV Cache for Coordinated Waiting and Execution Time in LLM Serving

Large Language Model (LLM) serving systems face a KV-cache (KVC) bottleneck. In this paper, our experimental study shows that block-based allocation increases Time-Between-Tokens (TBT) due to preemptions, while prediction-based allocation increases Time-to-First-Token (TTFT) and TBT due to allocated but unused KVC and...

Haiying Shen, Tanmoy Sen, Masahiro Tanaka · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.