Skip to content

FestDPO: Few-step Generator Alignment with Direct Preference Optimization

Sep 2026 · 0 citations · 99 references
Computer Science

TL;DR

This work introduces Few-step DPO (FestDPO), an extension of DPO for few-step generative models that leverages nonparametric likelihood estimation from empirical samples that makes sample-based approximation of DPO loss computationally feasible.

Abstract

Few-step generative models can generate high-fidelity samples within a few function evaluations. Despite this efficiency, generated samples may not exhibit desirable properties. When these properties are difficult to encode as an explicit reward function, direct preference optimization (DPO) can align generative models using pairwise preference feedback without training a separate reward model. However, extending DPO to few-step generative models is challenging because few-step generative models are generally implicit, making the likelihood evaluation required by DPO intractable. To address this challenge, we introduce Few-step DPO (FestDPO), an extension of DPO for few-step generative models that leverages nonparametric likelihood estimation from empirical samples. By exploiting the fast sampling capabilities of few-step generative models, our approach makes sample-based approximation of DPO loss computationally feasible. Furthermore, the sample-based formulation makes FestDPO agnostic to the model family and sampling procedure. Our toy experiment demonstrates that FestDPO matches the reward-tilted target distribution across four few-step generators. For real-world tasks, we evaluate FestDPO in two domains: text-to-image generation and protein backbone generation. In text-to-image generation, FestDPO outperforms preference optimization baselines in both win rates against the base models and human evaluation scores. In protein backbone generation, it achieves a higher $\beta$-sheet fraction and better structural designability than the baselines.

View source

Similar papers

#machine learning Preprint Sep 2026

Aligning One-Step Generative Models with Reward-Weighted Transport Distillation

Theoretical analysis shows that the fixed-point distributions of RWTD interpolate between off-policy reward tilting of the reference and on-policy tilting of the current model, providing a principled approach to balancing reward adaptation with retention of prior knowledge.

Austin S. Wang, Zi-Heng Cheng, Le-Xing Ying · 0 citations
#machine learning Preprint Aug 2026

Reward-guided Fine-Tuning of One-Step Generative Models via Wasserstein Gradient Flow

This work considers one-step generators from an optimal transport view, investigating Wasserstein Gradient Flow (WGF) for modeling smooth and controlled distributional evolution in probability space, and proposes a novel reward-guided fine-tuning of a one-step generative model via WGF.

Hoseong Hwang, Woorim Han, Joungin Chun et al. · 1 citation
#artificial intelligence Preprint Sep 2026

Direct Preference Density Alignment for Conversational Audio Equalization

Large Language Model alignment typically relies on learned proxy reward models, which significantly increase the memory footprint during training and are notoriously prone to instability and reward hacking. While offline methods like Direct Preference Optimization (DPO) bypass the reward model, they lose the ability to...

Ioannis Stylianou, S. Shepstone, Jon Francombe et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Manifold-Constrained Initial Noise Optimization for Efficient Generative Model Alignment

Recent advances in distillation and flow-map models have enabled deterministic one- or few-step generation for high-quality data, facilitating a new branch of reward alignment approaches that directly optimize the initial noise from a Gaussian distribution. However, most existing initial-noise optimization methods rely...

Jin-Ho Chang, Jong Chul Ye · 0 citations
Conference 2026

Shortcut Diffusion Training With Cumulative Consistency Loss: An Optimal Control View

This paper forms few-step generation as a controlled base generative process, and shows that self-consistency loss can be understood through the lens of optimal control, and draws a connection between this approach and reinforcement learning, potentially opening the door to a new set of approaches for few-step generati...

Paribesh Regmi, S. Ghimire, Rui Li · 0 citations
#computer vision Preprint Aug 2026

RL-Native Distillation: Exploiting Scored Trajectories for Few-Step Image Generation

This work proposes REST (Reward-Enhanced Scored-Trajectory Distillation), a single-stage RL-distillation co-training framework that attaches a decoupled student to an arbitrary RL teacher that enables few-step CFG-free inference that matches or surpasses its 40-step RL teacher, with an overall additional training cost...

Yuhan Li, Fan-Gao Zeng, Sicong Kang et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.