Fixed-Sequence Confidence-Interval Calibration for Finite-Sample Selective Risk Control
Selective prediction requires a statistically reliable rule for deciding which model outputs can be returned while controlling the error rate among accepted predictions. In large language model (LLM) applications, uncertainty scores provide useful ranking information, but directly thresholding such scores does not yiel...