Preprint
Aug 2026
ChainPrune: Evaluating and Reducing Redundancy in Long Chain-of-Thought Reasoning
This work proposes ChainPrune, a novel reasoning path semantic structural optimization method to efficiently and controllably synthesize self-generated high-quality training data and incorporates a DPO-based preference learning method combined with supervised loss, effectively mitigating false reward suppression.
Weihang Pan, Zhengxu Yu, Yuxiang Zhang et al.
· 1 citation