It is established that population-level validation alone is insufficient for equity assessment of digital health AI, motivating subgroup-disaggregated reporting as a default standard, and subgroup-disaggregated reporting as a default standard for personalized configurations.
Abstract
As CGM-based AI tools approach clinical deployment, whether their accuracy is equitable across patient demographics remains insufficiently tested. To enable this evaluation, we constructed FairGlucose, a 300-patient CGM cohort balanced across 12 demographic strata (age x gender x type 1/type 2 diabetes), with 132,480 forecasting samples and 3,945 unique behavioral events (meals, exercise, medication) logged by 81 patients. Benchmarking 33 models across four families on 2-hour glucose forecasting, we find that population-level external validation can conceal substantial subgroup disparities. Aggregate out-of-distribution metrics appear stable (approximately 1.0), yet subgroup-level ratios range from 0.8 to 1.4, with T1D patients showing 6 mg/dL higher prediction error than T2D (p<0.001). This disparity persists across all 33 models, suggesting a property of the prediction task rather than any single architecture. Further analysis shows that subgroup performance gaps align with the proportion of clinically hard cases, and that input-length sensitivity varies across demographics, motivating personalized configurations. Frontier LLMs underperform specialized neural models by 1-6 mg/dL; behavioral events contribute negligibly (approximately 0.1 mg/dL) even under oracle event access. These findings establish that population-level validation alone is insufficient for equity assessment of digital health AI, motivating subgroup-disaggregated reporting as a default standard.
Clinical risk models routinely achieve strong aggregate performance while producing materially different error rates across patient subgroups. Audit pipelines have been proposed to catch this, but their components are rarely stress-tested, so it is unclear which parts of an audit can be trusted and under what conditions. We present KAISEN, a five-phase audit pipeline covering subgroup stratification, disparity measurement, mechanism diagnostics, post-hoc mitigation, and drift monitoring, evaluated to the point of failure on a synthetic benchmark of 16 disease tasks, 15 social-determinant axes from Healthy People 2030, and three prespecified intersections. Four findings follow. (i) Significance tracks each axis's gap against its own minimum detectable effect: rank correlation between significance count and raw equalized-odds difference (EOD) across the 15 axes is rho = 0.56, rising to rho = 0.78 once EOD is standardized by that floor. (ii) Per-group threshold optimization reduces EOD in 48 of 48 held-out runs (paired delta = -0.285, 95% CI [-0.313, -0.252]), while group-wise Platt scaling -- the better calibrator -- behaves as a coin flip on EOD (19 of 48 runs improved, 95% CI [0.26, 0.55]) with mean effect near zero, so what an audit should report is the variance, not the average. (iii) The mechanism diagnostic classifies 144 of 144 controlled cases correctly but recovers none of 48 model-driven cases under proxy misspecification, with no signal that it failed. (iv) CUSUM failures and false alarms track cohort realization far more than disease: at the reference threshold, all 27 false alarms and 7 of 8 missed shifts come from different seeds (chi-squared p = 0.002), so a threshold tuned on one cohort fails to transfer. All results are synthetic with known ground truth and do not establish clinical validity. Code, artifacts, and scripts reproducing every number are released.
Sparsh Roy, Samuel Girmachew, Nishita Chavan· 0 citations
Fairness audits of clinical AI models rarely make the evidentiary status of subgroup findings explicit: reassuring results may reflect insufficient statistical precision rather than true parity, and audit verdicts can easily reverse under equally defensible analytic choices. We introduce an evidence classification scheme that screens for sample size and precision, and integrates stability across design alternatives directly into the fairness claim. We demonstrate this scheme on the estimation of the brain-age gap (BAG), a potential clinical biomarker, from structural MRI using the Alzheimer's Disease Neuroimaging Initiative (ADNI) data. The male-female and Black-vs-White differences, along with the White-Male and Black-Female intersectional contrasts, are all classified as equivalence supported, stable across regressor choice (ridge vs. gradient-boosted trees) and feature representation (full feature set vs. cortical-thickness-only). The Asian-vs-White and Black-Male comparisons remain classified as insufficient data throughout, as neither meets the pre-specified minimum-sample threshold. The proposed scheme provides a path from raw fairness findings to justified fairness claims via pre-specified thresholds, minimum-information screening, and stability checks across declared design choices.
D. Stark, K. Ritter, Alzheimer's Disease Neuroimaging Initiative· medRxiv· 0 citations
Predictive models employing artificial intelligence (AI) and machine learning (ML) are increasingly being used for decision support in healthcare settings. These models may exhibit differential performance across population subgroups defined by race, age, sex, and other factors and cause disparate clinical impacts, leading to intensive recent study of what has been termed"model fairness". While many methods have been proposed to assess risk prediction model fairness, these techniques generally require that the end user pre-specify the groups across which fairness is to be evaluated. In real-world settings, however, important model performance disparities may arise in unknown subgroups defined by multiple intersecting characteristics. To address this problem, we propose the unfairness tree (utree), a data-driven recursive partitioning framework for identifying subgroups with differential model performance. In simulations, the utree exhibits nominal empirical type I error rates and good ability to detect, quantify, and characterize performance discrepancies defined by higher-order variable interactions. In six mortality risk models fit to the GUSTO-I acute myocardial infarction trial dataset, utrees identified subgroup-specific performance patterns, with age, sex, blood pressure, and Killip class consistently associated with differential model performance.
Abstract Objectives To develop a generalizable framework for identifying algorithmic discrimination risks arising from subgroup imbalances in machine learning training data, with relevance to medical informatics applications where heterogeneous real-world data can bias model behavior. Materials and Methods We introduce a discrimination risk assessment framework for training datasets, a 4-step methodology integrating: (1) controlled representation sampling, (2) ensemble-based model training, (3) multilevel subgroup disparity quantification, and (4) mitigation-oriented interpretation. The framework systematically perturbs training subgroup composition while holding evaluation sets fixed to isolate representation effects. Validation was performed across 7 publicly available pediatric type 1 diabetes datasets using sociodemographic variables and continuous glucose monitoring data to assess robustness under heterogeneous data sources. Results Representation balance alone does not guarantee stable or equitable model outputs. Ensemble analyses revealed subgroup-dependent volatility, with some groups consistently contributing to model generalization, while others inducing instability despite increased representation. These findings demonstrate structural sensitivity to data composition that is not captured by standard performance-only evaluations. Discussion The framework exposes disparities and instability patterns that remain hidden under single-model or balanced-data evaluations. By quantifying representation-driven behaviors and subgroup-specific training value, it offers a diagnostic tool for understanding data-induced risks prior to model deployment. Conclusion This work provides a methodology for assessing discrimination risks in training datasets used in medical informatics. By integrating data composition effects, prediction stability, and subgroup-level disparity analysis, the framework supports more reliable and transparent development of machine learning systems across diverse clinical and nonclinical contexts. Clinical trial registration This research did not involve any new clinical trial. All analyses were performed on data from previously registered clinical trials, as listed below: PEDAP dataset: The source of the data is the PEDAP Trial Study Group. (2024). PEDAP Public Dataset (Release 4). Retrieved from https://public.jaeb.org/dataset/599. The analyses content and conclusions presented herein are solely the responsibility of the authors and have not been reviewed or approved by PEDAP Trial Study Group. ClinicalTrials.gov Identifier: NCT04796779; registered March 15, 2021. IOBP2 RCT dataset: The source of the data is the Insulin Only Bionic Pancreas Pivotal Trial. Retrieved from https://public.jaeb.org/dataset/579. The analyses content and conclusions presented herein are solely the responsibility of the authors and have not been reviewed or approved by the Bionic Pancreas Research Group or Beta Bionics. ClinicalTrials.gov Identifier: NCT04200313; registered December 16, 2019. DCLP5 dataset: The source of the data is the University of Virginia. DCLP5 Dataset. Retrieved from http://public.jaeb.org/dataset/535. The analyses content and conclusions presented herein are solely the responsibility of the authors and have not been reviewed or approved by the University of Virginia. Associated ClinicalTrials.gov Identifier: NCT03844789; registered February 18th, 2019 DCLP3 dataset: The source of the data is the University of Virginia. The International Diabetes Closed Loop (iDCL) trial. DCLP3 Public Dataset (Release 3). Retrieved from http:///public.jaeb.org/dataset/573. The analyses content and conclusions presented herein are solely the responsibility of the authors and have not been reviewed or approved by the University of Virginia. ClinicalTrials.gov Identifier: NCT03591354; registered July 19, 2018. CITY dataset: The source of the data is the Jaeb Center of Health Research. CGM Intervention in Teens and Young Adults with T1D (CITY Public Dataset). Retrieved from http://public.jaeb.org/dataset/565. The analyses content and conclusions presented herein are solely the responsibility of the authors and have not been reviewed or approved by the Jaeb Center of Health Research. ClinicalTrials.gov Identifier: NCT03263494; registered August 28, 2017. SENCE public dataset: The source of the data is the Jaeb Center of Health Research. SENCE Public Dataset. Retrieved from http://public.jaeb.org/dataset/554. The analyses content and conclusions presented herein are solely the responsibility of the authors and have not been reviewed or approved by the Jaeb Center of Health Research. ClinicalTrials.gov Identifier: NCT02912728; registered September 23, 2016.
I. Bilionis, Ricardo C. Berrios, A. de Arriba Muñoz et al.· JAMIA Open· 0 citations
Abstract Background Fairness evaluation is essential for trustworthy clinical risk prediction. However, existing fairness-oriented discrimination metrics either ignore cross-group comparisons or rely on exhaustive pairwise evaluations, making them difficult to interpret and impractical for model selection. Objective This study aimed to develop and evaluate novel fairness-oriented discrimination metrics for clinical risk prediction that address limitations of within-group and pairwise cross-group approaches. Methods We examined theoretical properties of existing U-statistic–based metrics, including concordance index (CI) and area under the receiver operating characteristic curve (AUC), when applied to subgroups. We highlighted the distinction between within-group discrimination (ranking within a subgroup) and group-level discrimination (ranking relative to the broader population). Building on this framework, we proposed group-level extensions of the CI and AUC that summarize subgroup-specific performance in a single interpretable measure. We then applied these metrics to the PREVENT (Predicting Risk of Cardiovascular Disease Events) equation, a recently developed model for atherosclerotic cardiovascular disease. Results The traditional subgroup-specific CI and AUC captured within-group but not group-level discrimination, obscuring inequities in clinical decision-making. Existing cross-group approaches (eg, the xCI and xAUC metrics) addressed this limitation but became computationally and interpretively burdensome with multiple subgroups due to pairwise comparisons. Our proposed metrics provided a streamlined alternative, yielding 1 summary statistic per subgroup while retaining sensitivity to cross-group ranking disparities. Applied to PREVENT, these metrics revealed differences in subgroup performance not apparent from within-group evaluations. Conclusions By distinguishing between within-group and group-level discrimination, our framework clarifies a common source of misinterpretation in fairness evaluation. The proposed group-level extensions of the CI and AUC provide practical, interpretable tools for evaluating fairness in clinical prediction models, enabling more transparent and equitable risk assessment.
Haoyuan Wang, Chuan Hong, Michael J. Pencina et al.· JMIR AI· 0 citations