Calibration-Aware Reinforcement Learning for Large Language Models: A Survey of Objectives, Optimization, and Decision-Making
Abstract
Large language models increasingly emit confidence reports, predictive distributions, and typed decisions that determine whether a system answers, abstains, retrieves evidence, or spends more computation. We survey calibration-aware reinforcement learning (RL), in which a reported probability is scored by the reward, consumed by the policy’s actions, or both. Such a probability means something only relative to an event, the reporter’s information, and a population or policy. Unlike post-hoc calibration, RL can change the answers being assessed, the incentive actually optimized, and the behavior that consumes the number. We organize the literature around three gaps. Between reporting and capability, proper-score geometry shows that a joint answer-confidence reward decomposes into terms for mean accuracy, the distribution of success across inputs, and reporting error, so it can prefer a less accurate policy even under truthful reporting. Between the stated objective and the implemented update, group standardization, the treatment of parameter-dependent rewards, and finite-ensemble scoring can change what training optimizes. Between ranking and decision value, a score can keep its ordering while losing the numerical meaning that a cost-derived threshold requires. Exact constructions make each gap concrete, and a method landscape traces representative protocols to primary sources. We then give a comparator guide; an evaluation contract that separates reporting gains from changes in the answer policy, effects of the implemented update, and operating-point selection; and open questions. The practical conclusion is a comparator rule: post-hoc fitting is the comparator to beat when predictions can remain fixed, and direct supervision when the desired report has a tractable differentiable loss; RL is a distinct intervention only when reasoning, retrieval, answering, abstention, or information acquisition must adapt through outcome feedback.