Skip to content

Author

Vishakha Ramani

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

An Approximate Queueing Model of LLM Inference Serving for SLO-Driven Autoscaling

Performance models of LLM servers support both latency evaluation and the design of controllers for autoscaling against service level objectives (SLOs) and for inference optimization. We model the multiplexed execution of prefill and decode operations with a tractable, approximate queueing model under Markovian assumpt...

Vishakha Ramani, Asser N. Tantawi · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.