Skip to content
Open access

Evaluating Uncertainty Quantification in Clinical Machine Learning: Calibration, Robustness, and Decision Utility under Distribution Shift

2026 · International journal of research and innovation in applied science · Vol 11, pp. 1241-1252 · 0 citations

TL;DR

A rigorous empirical framework is presented for comparing three uncertainty quantification approaches on two clinical prediction tasks, in-hospital mortality and 30-day readmission, using 74,829 ICU admissions from the MIMIC-IV database to support a more demanding evaluation standard for UQ in clinical machine learning.

Abstract

Machine learning models deployed in clinical decision support systems almost universally produce point predictions without any accompanying measure of uncertainty. In high-stakes healthcare settings, this is not merely a technical limitation: a miscalibrated prediction can directly influence patient management decisions with real consequences for safety and outcomes. Several uncertainty quantification (UQ) methods have been proposed to address this gap, including conformal prediction, Bayesian neural networks (BNNs), and Monte Carlo (MC) dropout; however, their comparative evaluation has predominantly been conducted under idealised conditions that do not reflect clinical deployment, where patient populations, treatment practices, and data recording procedures change over time. We present a rigorous empirical framework for comparing these three UQ approaches on two clinical prediction tasks, in-hospital mortality and 30-day readmission, using 74,829 ICU admissions from the MIMIC-IV database. Methods are assessed across three dimensions: calibration quality (ECE, ACE, Brier score), robustness under temporal distribution shift, and clinical decision utility via net benefit analysis. All experiments are replicated across five independent seeds, with comparisons made using Wilcoxon signed-rank tests with Holm-Bonferroni correction. Under standard evaluation conditions, all three methods achieve similar discriminative performance (AUROC 0.836-0.844 for mortality; 0.637-0.641 for readmission). Under temporal shift, BNN calibration degrades most sharply on the readmission task (ΔECE = 0.011 ± 0.002) compared with MC Dropout (ΔECE = 0.002 ± 0.003), while AUROC paradoxically improves for all methods, demonstrating that discriminative and calibration performance can decouple under distribution shift. Conformal prediction maintains near-nominal empirical coverage on the mortality task (0.886 ± 0.002) but shows notable violations on readmission, raising practical concerns about exchangeability assumptions in deployed systems. These findings support a more demanding evaluation standard for UQ in clinical machine learning, one that moves beyond static i.i.d. benchmarks toward temporally robust, decision-aware assessment.

Read PDF

Similar papers

Preprint Jul 2026

Bayesian uncertainty estimation improves clinical decision making in medical AI agents

Machine learning models for medical image analysis typically lack a reliable measure of confidence, limiting their use in ambiguous or atypical cases. Here we show that Monte Carlo dropout, applied to a multi-task chest-radiograph classifier (eight thoracic findings, 137,593 training images), provides an epistemic uncertainty signal that tracks generalisation across training-set scales and flags confident yet error-prone predictions. Adding this signal to the point prediction raised error-detection AUROC from 0.74 to 0.77 ($\Delta$AUROC +0.023, 95% CI [+0.014, +0.033]). In a controlled 2x2 factorial experiment, a clinical-decision-support agent exploited this uncertainty only when it was delivered as a binary error-risk flag rather than as raw scores, cutting confident misdiagnoses on unreliable findings from 8.5% to 2.7%. Epistemic uncertainty estimation thus carries decision-relevant information beyond point predictions, but its value for downstream agents depends on how it is communicated.

Frederik Hauke, P. Wienholt, Christiane Kuhl et al. · 0 citations
Open access Jul 2026

Translational Evaluation of Interpretable Machine Learning for Cardiovascular Risk Prediction: Calibration Decomposition, Subgroup Audits, and Decision-Utility Analysis

Cardiovascular disease risk prediction models are often evaluated primarily by discrimination, although translational decision making depends on well-calibrated probabilities, subgroup reliability, and demonstrated clinical utility at actionable risk thresholds. This study conducted a translational evaluation of interpretable machine-learning models for heart disease prediction using a deployment-oriented framework integrating discrimination, calibration (including Murphy decomposition), explainability, subgroup stability, and decision-utility analysis via decision curve analysis. Using a large secondary dataset (308,774 observations; 19 predictors; prevalence 8.1%), models were trained with a stratified hold-out design and evaluated on a fixed test set. Histogram-based gradient boosting achieved the strongest discrimination (PR-AUC 0.3177; AUROC 0.8407) and strong probabilistic accuracy (Brier score 0.0633; ECE 0.0045), with Murphy decomposition indicating minimal reliability loss while preserving resolution. Explainability analyses (SHAP with PDP/ALE/ICE diagnostics) enabled transparent assessment of feature contributions and nonlinear effects relevant to plausibility and governance. Subgroup analyses indicated broadly stable discrimination but more variable calibration across age and self-reported general health strata, supporting the need for subgroup-aware monitoring. Decision curve analysis demonstrated positive net benefit relative to treat-all and treat-none strategies across screening-relevant thresholds (0.05–0.15), with workload trade-offs informing threshold selection for practice.

Hathaichanok Chompoopong, Warawut Narkbunnum · 0 citations
Preprint Aug 2026

A Statistical Framework for Data-Driven Discovery of Differential Performance in Clinical Risk Prediction Models

Predictive models employing artificial intelligence (AI) and machine learning (ML) are increasingly being used for decision support in healthcare settings. These models may exhibit differential performance across population subgroups defined by race, age, sex, and other factors and cause disparate clinical impacts, leading to intensive recent study of what has been termed"model fairness". While many methods have been proposed to assess risk prediction model fairness, these techniques generally require that the end user pre-specify the groups across which fairness is to be evaluated. In real-world settings, however, important model performance disparities may arise in unknown subgroups defined by multiple intersecting characteristics. To address this problem, we propose the unfairness tree (utree), a data-driven recursive partitioning framework for identifying subgroups with differential model performance. In simulations, the utree exhibits nominal empirical type I error rates and good ability to detect, quantify, and characterize performance discrepancies defined by higher-order variable interactions. In six mortality risk models fit to the GUSTO-I acute myocardial infarction trial dataset, utrees identified subgroup-specific performance patterns, with age, sex, blood pressure, and Killip class consistently associated with differential model performance.

A. Neher, Julian Wolfson · 0 citations
Review Open access Jul 2026

Advancements in reinforcement learning for clinical decision-making in healthcare: a systematic review

Clinical decision-making increasingly relies on data-driven tools, but most systems today are still predictive models that work at isolated time points. Reinforcement learning (RL) provides a different approach by optimizing sequences of actions under uncertainty. It’s often seen as a foundation for more “agentic”AI in healthcare. We conducted a systematic literature review of RL-based clinical decision support systems (CDSS) published between 2020 and January 2026. We reviewed 66 studies, looking at the clinical domain, decision type, RL methods, data, and system maturity. RL-based CDSS are mostly used in critical care, cardiology, oncology, and diabetes, focusing on therapeutic dosing optimization. Actor-critic and policy-gradient methods are mainly used in continuous physiological/device-control settings. Most systems are trained offline using historical data: 66.7% rely on observational clinical data, 21.2% use simulated environments, and 12.1% combine both. Overall, RL-based CDSS are still partially autonomous, they often prioritize autonomy and personalization over runtime oversight, interpretability, and evaluation. We suggest an “agentic readiness”framework to address these gaps and emphasize the need for better safeguards, clearer reporting, and more human-centered assessments.

Ali Najafi, Amirfarhad Farhadi, A. Zamanifar · 0 citations
Review Open access Aug 2026

Explainable AI in Healthcare: A Comparative Analysis of Interpretability Techniques for Clinical Decision Support Systems

Artificial intelligence has made a great impact on healthcare by providing accurate disease diagnosis, personalised treatment regimens, and efficient clinical decision making. But many of the advanced machine learning and deep learning models are black-box systems, and healthcare professionals find it difficult to understand the logic behind their predictions. This opacity hinders the adoption of intelligent systems in clinical settings where trust and accountability are a must. In this review paper we compare the main interpretability techniques that have been used in clinical decision support systems. These techniques include Local Interpretable Model-Agnostic Explanations (LIME), SHapley Additive exPlanations (SHAP), saliency maps, Gradient-weighted Class Activation Mapping (Grad-CAM), attention mechanisms, and decision trees, among others. We performed a systematic literature review to evaluate these techniques based on interpretability, computational complexity, scalability, transparency, and clinical relevance. A systematic literature review was performed to evaluate the techniques in terms of interpretability, computational complexity, scalability, transparency and clinical relevance. The analysis shows that SHAP provides complete local and global explanations, while LIME provides computationally efficient local interpretations. Visualisation based methods such as Grad-CAM and saliency maps are especially useful for medical image analysis, while attention mechanisms are suitable for sequential healthcare data. The study concludes that explainable artificial intelligence improves trust, reliability, and accountability in healthcare systems and is a prerequisite for successful integration of intelligent technologies into clinical practice. Keywords: machine learning; clinical decision support systems; Explainable Artificial Intelligence; Healthcare Analytics; interpretability

Riya Jacob K · 0 citations
Conference Jul 2026

Bridging Predictive Modelling and Healthcare Workflows: A Machine Learning-Enabled Clinical Decision Support System

Clinical decision support systems (CDSS) have demonstrated the potential to improve healthcare quality by delivering timely, patient-specific insights at the point of care; however, many machine learning (ML)–based approaches remain confined to research settings due to challenges in system integration, workflow alignment, and deployability. This paper presents the design and evaluation of a ML-enabled CDSS that bridges comparative predictive modelling with a modular system architecture suitable for clinical deployment. Multiple supervised ML models, including Random Forests, Logistic Regression, Support Vector Machines, Neural Networks, and XGBoost, were trained and evaluated to predict future serum potassium levels using routinely collected laboratory biomarkers. Model performance was assessed using regression-based metrics, including mean absolute error (MAE), mean absolute percentage error (MAPE), root mean squared error (RMSE), and coefficient of determination (R^2). Evaluation using de-identified data from patients in British Columbia, Canada, demonstrated that the Random Trees model achieved the best overall regression performance (MAE = 0.28, RMSE = 0.37, R^2= 0.61, MAPE = 6.32). This model was selected for deployment within the proposed Predictive Risk Indicator for Serum Potassium Measurement (PRISM) clinical decision support system, enabling real-time generation of individualized predictions mapped to clinically meaningful risk categories. This work demonstrates the feasibility of integrating comparative ML model development with system-level integration through a model-as-a-service paradigm. The proposed system provides a foundation for future prospective validation, usability evaluation, and integration into routine clinical workflows.

K. Renganathan, Waqar Haque, Anurag Singh et al. · 0 citations