Towards Bridging the Gap Between Offline and Iterative Alignment via Preference Distillation
This work proposes Distilled Preference Probability Policy Optimization (DP3O), an effective and efficient offline alignment algorithm that outperforms state-of-the-art offline methods, achieves performance comparable to iterative DPO, and reduces training time, demonstrating both its effectiveness and efficiency.