Entropy-Valley (EV), a training-free length selector that scores candidate target canvases by mean predictive entropy from all-mask forward passes and selects the canvas the backbone is most prepared to fill, is introduced.
Abstract
Machine translation tests masked diffusion language models (dLLMs) because every source token must be rendered faithfully, while fixed canvas decoding must choose target length before denoising. Existing masked diffusion decoding work mainly studies token unmasking order, leaving this length decision under-explored despite its direct effect on coverage and redundancy. We introduce Entropy-Valley (EV), a training-free length selector that scores candidate target canvases by mean predictive entropy from all-mask forward passes and selects the canvas the backbone is most prepared to fill. Relative to a baseline using training corpus length statistics, EV recovers 64.9%, 65.3%, and 33.0% of the COMET-22 gain from reference target lengths on En$\to$Zh, Zh$\to$En, and En$\to$De. Our diagnostics show that denoising-friendly lengths need not match reference lengths. Evaluation by three translation experts supports the En$\leftrightarrow$Zh adequacy gains, with stronger evidence on Zh$\to$En. Compared with a LLaMA-3-8B autoregressive (AR) model trained on the same fine-tuning data, the EV system ties on En$\to$Zh and leads on Zh$\to$En; an oracle-length diagnostic further shows that, in this masked diffusion MT setting, deciding which tokens to reveal first matters less than how the target length is supplied.
CARVE (Counterfactual-Aware Reveal with Verified Expansion), a training-free variable-length algorithm for masked diffusion LMs that consistently improves average performance over fixed-length baselines across all evaluated model families.
Wail Bouhedja, Amr Mohamed, Guo-Kan Shang· 0 citations
DARD is proposed, a training-free framework that separates tokens into masked, candidate, and unmasked states and adaptively regulates their influence on subsequent decoding, and consistently improves the speed-quality Pareto frontier over recent revocable decoding methods.
Woo-Soon Park, Insu Lee, Minyoung Noh et al.· 1 citation
Masked diffusion language models (MDLMs) can generate text efficiently by predicting multiple masked tokens in parallel, but predictions from the same forward pass are not necessarily reliable when committed together. We study when parallel commitment is reliable. Our diagnostics show that confidence alone does not det...
Zheng-Hao He, Bo-Han Liu, Guang-Zhi Xiong et al.· 0 citations
Diffusion language models (DLMs) enable parallel generation by predicting and committing multiple tokens at each denoising step, yet they can generate individually plausible but mutually inconsistent tokens. Recent work shows that \emph{soft tokens} can mitigate this issue by representing uncertain positions with conti...
A black-box, inference-time diagnostic that tells these two cases apart without retraining or annotation is introduced, with the pattern holding on English-Marathi and English-Tamil, with the failure modes tracking the post-edit distribution rather than MT quality or language family.
Isuru Wijesiri, Nisansa de Silva, Kavindu Warnakulasuriya et al.· 0 citations
A language-level router based on paired e-processes that continuously compares direct and translation-assisted classification before freezing a routing policy is introduced, demonstrating that paired e-processes enable statistically controlled, anytime-valid, and auditable multilingual classification routing.
Wajdi Ben Saad, Safa Madiouni· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.