Preprint
Jul 2026
Sangam: Efficiently Serving Diffusion LLMs with the AR Stack
Sangam, a serving system for cached dLLM inference that adopts a hybrid serving strategy, overflowing prefills onto decode workers to relieve prefill under-provisioning, and uses the same deficit-budget scheduler to protect those workers'decodes from the overflow.
Nitin Kedia, Saurabh Agarwal, Myungjin Lee et al.
· 1 citation