Calibration-Aware Reinforcement Learning for Large Language Models: A Survey of Objectives, Optimization, and Decision-Making
Abstract
Large language models increasingly produce confidence reports, predictive distributions, and structured decisions that determine whether a system answers, abstains, retrieves evidence, or spends additional computation. Reinforcement learning can improve these signals, but it can also change the answers being assessed, the meaning of the training target, and the effective scoring incentive. We survey calibration-aware reinforcement learning through the connection between probability targets, objectives, implemented updates, and decision use. The synthesis combines a comparison of representative methods, with claims traced to primary sources, and three mechanisms for interpreting divergent results. Classical proper-score geometry separates mean capability, the distribution of conditional success probabilities, and reporting regret in joint answer--confidence training. Group transformations, parameter-dependent rewards, and finite-ensemble scoring distinguish an ideal population objective from the update actually implemented. Cost-sensitive decisions distinguish useful ordering from probability magnitudes that can be reused across operating points. These distinctions organize comparisons with post-hoc calibration, supervised reporting, selective ranking, direct action learning, and numerical distribution prediction. We provide a method landscape, an evaluation contract, and a research agenda for separating reporting improvements from policy changes, estimator effects, and threshold adaptation. The resulting perspective treats calibration as a probability claim whose conditions must remain explicit throughout learning and deployment, and identifies when RL supplies an intervention that simpler calibration methods do not.