Preprint
Aug 2026
Dependency-Aware Revocable Decoding for Efficient Diffusion Large Language Model Inference
DARD is proposed, a training-free framework that separates tokens into masked, candidate, and unmasked states and adaptively regulates their influence on subsequent decoding, and consistently improves the speed-quality Pareto frontier over recent revocable decoding methods.
Woo-Soon Park, Insu Lee, Minyoung Noh et al.
· 1 citation