Skip to content

Author

Nitin Kedia

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Jul 2026

Sangam: Efficiently Serving Diffusion LLMs with the AR Stack

Sangam, a serving system for cached dLLM inference that adopts a hybrid serving strategy, overflowing prefills onto decode workers to relieve prefill under-provisioning, and uses the same deficit-budget scheduler to protect those workers'decodes from the overflow.

Nitin Kedia, Saurabh Agarwal, Myungjin Lee et al. · 1 citation