Skip to content

Author

Aviral Kumar

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Jun 2026

Addressing Over-Refusal in LLMs with Competing Rewards

The resulting model SEAR deliberately engages in harmful reasoning as exploration while reliably flipping back to a safe answer, demonstrating that this behavior helps mitigate over-refusal and defend against attacks that directly manipulate the reasoning to be harmful.

Taeyoun Kim, Aviral Kumar · 0 citations