Skip to content
Open access

Fixed-Sequence Confidence-Interval Calibration for Finite-Sample Selective Risk Control

Oct 2026 · Stats · 0 citations

Abstract

Selective prediction requires a statistically reliable rule for deciding which model outputs can be returned while controlling the error rate among accepted predictions. In large language model (LLM) applications, uncertainty scores provide useful ranking information, but directly thresholding such scores does not yield finite-sample risk guarantees. We propose confidence-interval calibration (CIC), a general calibration framework that converts an arbitrary uncertainty score into a risk-controlled selective decision rule. Given a held-out calibration sample, CIC estimates the acceptance-conditioned error rate for a pre-specified sequence of candidate thresholds and constructs one-sided upper confidence bounds using either Hoeffding’s inequality or the exact Clopper–Pearson procedure. Candidate thresholds are examined sequentially, and CIC returns the most permissive threshold within the consecutively certified prefix. Under independent and identically distributed calibration and deployment samples, we establish a finite-sample, high-probability guarantee that any returned threshold controls the acceptance-conditioned error rate at the prescribed level. In our experiments, calibration samples on the order of 103 examples already provide useful certification in favorable regimes, although the required sample size depends on the target risk level, model error rate, and quality of the uncertainty ranking. Experiments on CommonsenseQA and TriviaQA across seven LLMs demonstrate reliable empirical risk control with favorable acceptance rates across a range of target risk levels.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.