Skip to content

Author

Alexander Korotin

4 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#machine learning Preprint Oct 2026

One-Step Generation via Riemannian Wasserstein Gradient Flows

Recently, Drifting Models and Wasserstein Gradient Flows have attracted substantial attention because they move iterative distributional refinement to training and amortize it into a generator, enabling fast inference. However, existing formulations have been developed largely for continuous Euclidean domains, such as...

David Li, Chanhyuk Lee, Jaehoon Yoo et al. · 0 citations
#machine learning Preprint Oct 2026

IDRF: Inverse-Distilled Reward Fine-tuning of Masked Discrete Diffusion Models

Masked discrete diffusion models offer a promising alternative to autoregressive generation, but iterative sampling can be costly, and intractable sequence likelihoods complicate reward fine-tuning. We introduce IDRF, a framework for reward fine-tuning of few-step masked discrete diffusion generators. Starting from a s...

V. Gromadskii, David Li, Samson Gourevitch et al. · 0 citations
#natural language process... Preprint Sep 2026

E-MoE: Enhanced Mixture-of-Experts for Non-Factorized Diffusion Language Models

Masked diffusion models (MDMs) generate sequences by progressively unmasking several tokens per denoising step, but their reverse process is typically factorized over positions, limiting sample quality in the few-step regime where diffusion's speed advantage over autoregressive decoding matters most. A recent line of w...

Arseny Ivanov, A. Kolesov, Alexander Korotin et al. · 0 citations
#machine learning Preprint Sep 2026

Alpha Diffusion Language Models: Factorization Alone Is Not the Problem

Discrete diffusion language models can generate multiple tokens in parallel, but reducing the number of denoising steps can lead to inconsistent predictions. Standard cross-entropy training fits conditional token marginals, whereas parallel generation requires consistent joint predictions. We introduce Alpha Diffusion...

N. Gushchin, Dmitry Baranchuk, Alexander Korotin · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.