Shared Experience, Separate Learning: Companion Confidence Calibration for LLMs
CoCal is introduced, which trains a lightweight companion from rollout hidden states and verifier-derived correctness supervision while leaving task optimization unchanged, and shows that CoCal improves confidence estimation without sacrificing task performance, outperforming both RL-based concurrent methods and matche...