#machine learning
Oct 2025
Large Reasoning Models Learn Better Alignment from Flawed Thinking
RECAP (Robust Safety Alignment via Counter-Aligned Prefilling), a principled reinforcement learning (RL) method for post-training that explicitly teaches models to override flawed reasoning trajectories and reroute to safe and helpful responses, substantially improves safety and jailbreak robustness, reduces overrefusal, and preserves core reasoning capability.
Sheng-Hsuan Peng, E. Smith, Ivan Evtimov et al.
· arXiv.org · 10 citations
· ⚡2