Large language models increasingly emit confidence reports, predictive distributions, and typed decisions that determine whether a system answers, abstains, retrieves evidence, or spends more computation. We survey calibration-aware reinforcement learning (RL), in which a reported probability is scored by the reward, c...
Large language models increasingly emit confidence reports, predictive distributions, and typed decisions that determine whether a system answers, abstains, retrieves evidence, or spends more computation. We survey calibration-aware reinforcement learning (RL), in which a reported probability is scored by the reward, c...
Large language models increasingly produce confidence reports, predictive distributions, and structured decisions that determine whether a system answers, abstains, retrieves evidence, or spends additional computation. Reinforcement learning can improve these signals, but it can also change the answers being assessed,...