Skip to content

Author

Ramesh Nampelly

We have 4 of 4 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Sep 2026

The Inference Engineering Pareto Atlas: Which Optimizations Dominate the Cost, Quality, and Latency Frontier?

LLM inference optimizations report speedups on different models, GPUs, prompts, and quality metrics, making them hard to compare or combine. We build a cost, quality, and latency Pareto atlas to identify the best configurations for different deployment constraints. Since exhaustive testing is impractical, we measure 54...

Srikanta Datta Tumkur, Jay Iyer, Mehar Simhadri et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Where Should the KV Cache Live? Placement Policies Across GPU, CPU, and SSD for Long-Lived Sessions

This work compares recency, reuse frequency, predicted reuse, and an EWMA predictor with prefetch lookahead across chat, agent, and document question answering workloads, and concludes that reuse frequency performs best for agents and document question answering and prefetching does not justify its bandwidth cost.

Srikanta Datta Tumkur, Jay Iyer, Mehar Simhadri et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Calibrate, Then Route: A Measured Study of Learned Request Routing for Disaggregated LLM Serving

A router is studied that estimates the additional completion time on each instance using exact prompt length, predicted output length, post admission KV cache pressure, and SLO class to match the goodput of round robin using six GPUs instead of seven.

Srikanta Datta Tumkur, Jay Iyer, Mehar Simhadri et al. · 0 citations
Preprint Aug 2026

More GPUs or a Smaller Cache? Tensor Parallelism versus KV Compression for Memory-Bound LLM Serving

When an LLM serving deployment runs out of KVcache room, there are two well-established ways out. Tensor parallelism shards the weights and the KV cache across two, four, or eight devices, buying memory headroom at the price of an all-reduce on every layer and a hardware bill that grows with the device count. The algor...

Srikanta Datta Tumkur, Mehar Simhadri, Anshu Bansal et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.