LLM inference optimizations report speedups on different models, GPUs, prompts, and quality metrics, making them hard to compare or combine. We build a cost, quality, and latency Pareto atlas to identify the best configurations for different deployment constraints. Since exhaustive testing is impractical, we measure 54...
Srikanta Datta Tumkur, Jay Iyer, Mehar Simhadri et al.· 0 citations
This work compares recency, reuse frequency, predicted reuse, and an EWMA predictor with prefetch lookahead across chat, agent, and document question answering workloads, and concludes that reuse frequency performs best for agents and document question answering and prefetching does not justify its bandwidth cost.
Srikanta Datta Tumkur, Jay Iyer, Mehar Simhadri et al.· 0 citations
A router is studied that estimates the additional completion time on each instance using exact prompt length, predicted output length, post admission KV cache pressure, and SLO class to match the goodput of round robin using six GPUs instead of seven.
Srikanta Datta Tumkur, Jay Iyer, Mehar Simhadri et al.· 0 citations
When an LLM serving deployment runs out of KVcache room, there are two well-established ways out. Tensor parallelism shards the weights and the KV cache across two, four, or eight devices, buying memory headroom at the price of an all-reduce on every layer and a hardware bill that grows with the device count. The algor...