Preprint
Jul 2026
Adaptive Inference Batching using Policy Gradients
RL's advantage over engineered heuristics concentrates in combinatorial, multi-resource decisions rather than single-resource temporal scheduling, a practical distinction for deciding where learned policies justify their engineering cost in production inference infrastructure.
Ruslan Sharifullin
· 0 citations