Skip to content
Open access

A framework for fair and robust clinical risk prediction through collaborative learning and localized uncertainty quantification

TL;DR

A collaborative learning framework that treats demographic subgroups as clients and aggregates their model parameters using strategies that encode different assumptions about group contribution, and a novel difficulty decomposition framework, derived from the localized conformal classifier's calibration scores, distinguishes sources of uncertainty arising from the local neighborhood from those specific to an individual prediction.

Abstract

Clinical prediction models trained on heterogeneous populations may exhibit disparities in predictive performance across patient subgroups, potentially reflecting differences in data availability, feature completeness, and representation in training data. Existing fairness-aware approaches can introduce performance-fairness trade-offs and commonly optimize group-level disparities without explicitly accounting for their subgroup-specific sources, limiting both their effectiveness and their clinical interpretability. This dissertation makes two methodological contributions. The first is a collaborative learning framework that treats demographic subgroups as clients and aggregates their model parameters using strategies that encode different assumptions about group contribution. Rather than optimizing an explicit group-fairness constraint, the framework preserves and integrates subgroup-specific information to improve equity while largely maintaining subgroup predictive performance. The second contribution is the localized conformal classifier, a neighborhood-adaptive conformal prediction framework for instance-level uncertainty quantification. By calibrating prediction sets according to the density and composition of a patient's local neighborhood in the feature space, the method communicates both the model's prediction and the uncertainty associated with that individual prediction. This information is particularly consequential when predictions inform clinical decision-making. Both contributions are evaluated across three clinical prediction tasks: 30-day mortality in ICU patients with sepsis, treatment non-completion in patients with substance use disorder, and hospital readmission after surgical resection in patients with colorectal cancer. Compared with reweighting and constrained optimization, collaborative learning achieves more favorable fairness-performance trade-offs with less degradation in predictive performance, and its variants consistently appear on the Pareto frontier across all three prediction tasks. A novel difficulty decomposition framework, derived from the localized conformal classifier's calibration scores, distinguishes sources of uncertainty arising from the local neighborhood from those specific to an individual prediction. SHAP attribution reveals how collaborative learning and fairness interventions alter the feature-attribution patterns in the underlying prediction models. Surrogate regression trees trained on the localized conformal classifier's difficulty scores relate neighborhood and instance difficulty to clinical features observable at the point of care. Clinical features remain strongly associated with instance-level difficulty under collaborative learning, whereas these associations weaken substantially after reweighting and are absent in some analyses. Across all three clinical contexts, local outcome heterogeneity is the most consistent driver of neighborhood-level predictive difficulty, while proximity to the decision boundary is the principal driver of instance-level difficulty. Together, these contributions provide a framework for developing and evaluating clinical prediction models that are not only accurate, but also more equitable across patient groups and more informative about the reliability of individual predictions.

Read PDF

Similar papers

#artificial intelligence Preprint Aug 2026

FairGlucose: A CGM Fairness Benchmark Reveals Subgroup Disparities Hidden in Population-Level Validation

It is established that population-level validation alone is insufficient for equity assessment of digital health AI, motivating subgroup-disaggregated reporting as a default standard, and subgroup-disaggregated reporting as a default standard for personalized configurations.

Junjie Luo, Xuzhe Zhi, Rui Han et al. · 0 citations
Preprint Sep 2026

A statistical framework for identifying subgroup vulnerability to predictive multiplicity in clinical AI

AI models trained on the same data can disagree about patient risk, with disagreement potentially concentrated in clinically important subgroups. We propose V(S), a statistically grounded vulnerability index combining an observable lower-bound witness of model disagreement with clinical severity, and develop inference...

Enock Adu Bonsu · 0 citations
Sep 2026

FairMoE-Health: Fairness-Aware Mixture of Experts for Equitable Multimodal Clinical Prediction.

The Routing Disparity metric and multi-level debiasing framework introduced here generalize to MoE systems operating on demographically heterogeneous populations, providing both an audit tool and an architectural intervention for a bias mechanism that existing fairness methods leave unaddressed.

Xiaoyang Wang, Christopher C. Yang · 0 citations
Open access Aug 2026

A Simplified Metric to Streamline Between-Group Fairness Assessment for Predictive Models: Algorithm Development and Evaluation Study

Abstract Background Fairness evaluation is essential for trustworthy clinical risk prediction. However, existing fairness-oriented discrimination metrics either ignore cross-group comparisons or rely on exhaustive pairwise evaluations, making them difficult to interpret and impractical for model selection. Objective Th...

Hao-Yuan Wang, Chuan Hong, Michael J. Pencina et al. · 0 citations
Preprint Aug 2026

External Risk Prediction Informed Bayesian Survival Analysis

Prognostic factor evaluation and prediction model development are central to precision oncology, enabling patient risk stratification and individualized treatment selection. Unified predictions that synthesize information from existing models are valuable for comprehensive and consistent risk assessment. Many studies a...

Yena Jeon, Yunxiang Huang, H. Kim et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.