FestDPO: Few-step Generator Alignment with Direct Preference Optimization
This work introduces Few-step DPO (FestDPO), an extension of DPO for few-step generative models that leverages nonparametric likelihood estimation from empirical samples that makes sample-based approximation of DPO loss computationally feasible.
Jaewoo Lee, Kyuil Sim, Hyeongyu Kang et al.
· 0 citations