2025
VPO: Reasoning Preferences Optimization Based on V-Usable Information
This work proposes VPO, a negative gradient constraint method for human non-preference samples based on V -usable information, which can alleviate the squeezing effect of DPO, enhance alignment with the generation objective, and maintain the model’s ability to distinguish between preference and non-preference samples.
Zecheng Wang, Chunshan Li, Yupeng Zhang et al.
· Neural Information Processin... · 1 citation