We introduce a post-training method for diffusion language models (DLMs) that minimizes Maximum Mean Discrepancy (MMD) between generated and reference distributions in the feature space of a frozen pretrained DLM. To estimate MMD, we retain contextual features at individual token positions, obtaining multiple observati...
Ilya Drobyshevskiy, I. Sudakov, Maksim Semenov et al.· 0 citations
Recently, Drifting Models and Wasserstein Gradient Flows have attracted substantial attention because they move iterative distributional refinement to training and amortize it into a generator, enabling fast inference. However, existing formulations have been developed largely for continuous Euclidean domains, such as...
David Li, Chanhyuk Lee, Jaehoon Yoo et al.· 0 citations
Discrete diffusion language models can generate multiple tokens in parallel, but reducing the number of denoising steps can lead to inconsistent predictions. Standard cross-entropy training fits conditional token marginals, whereas parallel generation requires consistent joint predictions. We introduce Alpha Diffusion...
N. Gushchin, Dmitry Baranchuk, Alexander Korotin· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.