Preprint
Jul 2026
Answer-Conditioned Chains of Thought Degrade Verifiable-Reasoning Distillation in Large Language Models
This work shows that training a strong instruction-tuned reasoning model on its own answer-conditioned chains sharply lowers its verifiable-reasoning accuracy, and generates answer-blind data, because no correctness filter can see this damage in the data.
Jungseob Lee, Seungyoon Lee, Suhyune Son et al.
· 0 citations