Skip to content

Author

Chunshan Li

We have 1 of 5 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

2025

VPO: Reasoning Preferences Optimization Based on V-Usable Information

This work proposes VPO, a negative gradient constraint method for human non-preference samples based on V -usable information, which can alleviate the squeezing effect of DPO, enhance alignment with the generation objective, and maintain the model’s ability to distinguish between preference and non-preference samples.

Zecheng Wang, Chunshan Li, Yupeng Zhang et al. · 1 citation