REFLEX: Rethinking MoE Inference as Refinement-Aware Compute Allocation in Diffusion Language Models
REFLEX is proposed, a training-free method that keeps the default router unchanged while reorganizing expert computation around the evolving refinement process, and introduces a coarse-to-fine hierarchy for expert-budget allocation that aligns computation with block-relative refinement roles while using the Frontier-Pr...