SOWRC combines alert labels with randomized audits through Horvitz–Thompson losses and a martingale-mixture boundary and shows that population risk is non-identifiable when any silent region has zero labeling probability.
Abstract
Selective labels create a support failure for prediction along dependent stochastic processes: alert-triggered events are observed, whereas silent periods are usually unlabeled. We model this mechanism as predictable inclusion on a filtered probability space and show that population risk is non-identifiable when any silent region has zero labeling probability. Selective-observation weighted risk control (SOWRC) combines alert labels with randomized audits through Horvitz–Thompson losses and a martingale-mixture boundary. It provides finite-sample calibration-population control under arbitrary temporal dependence subject to predictable design choices, conditional ignorability, positivity, bounded losses, and deterministic design envelopes, together with a prospective guarantee under an externally certified deployment-drift envelope and explicit error allocation. Extensions cover anytime monitoring, multiple losses, adaptive budgets, and estimated propensities. Synthetic maintenance and financial studies, a complete-log replay on a real dependent sensor series with 100 audit-mask replications, and 4000 selection-level validation runs demonstrate support recovery and conservative probabilistic risk control on deterministic threshold grids.
Efficient adaptive inference reduces computational demands, but selecting among policies on the same calibration data without controlling errors simultaneously can compromise safety conclusions. Practical deployment also requires explicit reporting of infeasibility, conditional efficiency analysis, and certification ac...
Sooyoung Jang, Siyeon Park, Sungpil Woo et al.· IEEE Access· 0 citations
Selective prediction requires a statistically reliable rule for deciding which model outputs can be returned while controlling the error rate among accepted predictions. In large language model (LLM) applications, uncertainty scores provide useful ranking information, but directly thresholding such scores does not yiel...
Rare clinical outcomes pose a difficulty deeper than ordinary class imbalance: a penalized logistic model can return finite, stable-looking coefficients before the data support a reliable threshold decision. We formulate the accrual question as decision-targeted sequential certification. On a prespecified finite monito...
These results support a scoped monitoring strategy for similar tabular settings: confidence-derived scores are effective for pointwise screening, whereas group-aware explanation audits provide complementary evidence about stable but incorrect feature reliance.
This work presents a protocol-aware empirical assessment across three settings: a C-MAPSS degradation-risk proxy, normal-only training for anomalous-sound detection on MIMII, and BDG2 forecasting-residual diagnostics with synthetic target perturbations.
Performance is closest to the oracle under autoregressive, threshold, stochastic-volatility, heavy-tailed, and variance-break dynamics, while local trends and regime switching are more challenging.
G. Vercellino· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.