This work proposes Distilled Preference Probability Policy Optimization (DP3O), an effective and efficient offline alignment algorithm that outperforms state-of-the-art offline methods, achieves performance comparable to iterative DPO, and reduces training time, demonstrating both its effectiveness and efficiency.
Abstract
Direct preference optimization DPO is a promising offline approach for aligning large language models (LLMs) due to its simplicity, computational efficiency, and implicit modeling of human preferences. Interestingly, iterative extensions of DPO have achieved stronger performance on academic benchmarks, raising two key questions: (i) Why do iterative methods generally outperform offline ones? (ii) Can their advantages be incorporated into offline alignment? To answer the first question, our controlled experiments reveal that the explicit preference model, additionally introduced in the iterative procedure, is a key factor behind its superiority over offline methods. This insight leads us to answer the second question affirmatively and propose Distilled Preference Probability Policy Optimization (DP3O), an effective and efficient offline alignment algorithm. DP3O first learns an explicit preference model using a helper class of LLMs and then distills its knowledge into policy optimization. Theoretically, we show that explicit preference modeling admits better estimation error control than implicit formulations, and that DP3O achieves a tighter generalization bound than hard-label DPO through variance reduction. Empirically, we evaluate DP3O on a wide range of chat-based and downstream tasks and show that it outperforms state-of-the-art offline methods, achieves performance comparable to iterative DPO, and reduces training time by about $42\%$, demonstrating both its effectiveness and efficiency.
This paper proposes and analyzes Comparison-based Preference Optimization (ComPO), a zeroth-order alignment method based on comparison oracles, and establishes a convergence guarantee for its basic offline scheme under smoothness, gradient sparsity, and compatibility between the oracle and a latent objective.
Peter Chen, Xi Chen, Wo-Tao Yin et al.· 0 citations
Personalizing large language models (LLMs) requires aligning generation behavior with user-specific preferences rather than aggregate quality. While Direct Preference Optimization (DPO) provides a stable framework for preference learning, its effectiveness in personalized settings critically depends on how preference p...
Ruo-Ming Jin, Xin-Yu Li, Hao Zhou et al.· 0 citations
Large Language Model alignment typically relies on learned proxy reward models, which significantly increase the memory footprint during training and are notoriously prone to instability and reward hacking. While offline methods like Direct Preference Optimization (DPO) bypass the reward model, they lose the ability to...
Ioannis Stylianou, S. Shepstone, Jon Francombe et al.· 0 citations
This work introduces FlowCPO, an offline forward-KL objective that uses both preferred and dispreferred samples without online rollouts and shows under explicit regularity conditions that the forward-KL objective is bounded by a contrastive flow matching loss, yielding a tractable surrogate on fixed data.
Yan-Sen Han, Shengyi Liao, Peng Sun et al.· 0 citations
This work forms the problem as obtaining a Nash equilibrium of a two-player zero-sum game between policies, and proposes two algorithms: Best-of-Nash (BoN) and Nash Mirror Descent (NMD), which are proved to achieve a duality gap that matches the problem lower bound.
Hadi Hosseini, Debmalya Mandal, Duo-Han Zhang· 0 citations
This work proposes DOTA, a data selection framework that minimizes the cost of generating preference data, while still ensuring the quality of training, and proposes a theoretically grounded metric called Preference Perplexity (PFP) that enables it to design a low cost, gradient-based method to effectively estimate the...
Chi Zhang, Jia-Chen T. Wang, Kun He et al.· Proceedings of the VLDB Endo...· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduAug 27, 2026
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.