When Do Models Admit They Are Wrong? Failure Disclosure Is Unstable Under Reinforcement Learning
Outcome-based reinforcement learning can produce models with similar task performance but very different ways of communicating about their mistakes. We study failure disclosure: whether a model admits that an attempted solution failed rather than staying silent or presenting it as successful. Across repeated outcome-on...