Compute-Adaptive Risk Control: A Finite-Sample Calibration Layer for Early-Exit and Cascade Inference
Abstract
Efficient adaptive inference reduces computational demands, but selecting among policies on the same calibration data without controlling errors simultaneously can compromise safety conclusions. Practical deployment also requires explicit reporting of infeasibility, conditional efficiency analysis, and certification across diverse policy families. We develop compute-adaptive risk control, a finite-sample calibration layer that treats policy selection as minimizing declared cost subject to a certified risk constraint: it returns the lowest declared-cost candidate certified within pre-specified chains or reports that none can be certified. The method combines bounded-loss tests with simultaneous error control and is evaluated through simulations, fixed-pool studies, and direct measurements on one workstation. In a non-clinical diagnostic involving 285 cases across 1,000 splits, at a 0.20 conditional-error target and 0.10 failure allowance, the multi-chain procedure returns a candidate in every split, with a held-out exceedance frequency of 0.049 and a mean relative cost of 1.541, compared with 4.510 for the scalar chain. In the primary image-benchmark setting, no certified policy is returned in any of 1,000 attempts. Three seed-distinct transformer classifier runs return policies in all 1,000 attempts, with unsafe-output frequencies from 0 to 0.003; both language-model configurations likewise return policies in all 1,000 attempts with zero unsafe outputs. Together, these results position explicit certification and justified withholding as a practical foundation for risk-aware, compute-efficient adaptive inference across diverse model families.