Jun 2026
Addressing Over-Refusal in LLMs with Competing Rewards
The resulting model SEAR deliberately engages in harmful reasoning as exploration while reliably flipping back to a safe answer, demonstrating that this behavior helps mitigate over-refusal and defend against attacks that directly manipulate the reasoning to be harmful.
Taeyoun Kim, Aviral Kumar
· arXiv.org · 0 citations