Fairness Audit and Debiasing of Machine Learning Models for Sepsis Mortality Prediction Across Demographic Subgroups: A Multi-Metric Study with Accuracy–Fairness Tradeoff Analysis
Aug 2026· International Journal of Innovative Science and Research Technology· 0 citations· 19 references
TL;DR
Three ML models were trained on a 10,000-patient cohort calibrated to published MIMIC-IV sepsis statistics, and a four-metric fairness audit was performed across race/ethnicity, sex, and insurance type.
Abstract
Machine learning (ML) models for sepsis mortality prediction are increasingly deployed in intensive care units,
yet performance across demographic subgroups remains poorly evaluated, and algorithmic disparities can directly
exacerbate health inequities in critical care settings. We trained three ML models—logistic regression (LR), XGBoost, and
a multilayer perceptron (MLP)—on a 10,000-patient cohort calibrated to published MIMIC-IV sepsis statistics, and
performed a four-metric fairness audit (equalized odds difference [EOD], demographic parity difference, predictive parity
gap, and subgroup calibration error) across race/ethnicity, sex, and insurance type. Per-group threshold optimisation was
applied for debiasing, and an accuracy–fairness Pareto tradeoff was quantified.
Fairness conclusions in clinical prediction can depend strongly on both the metrics reported and the demographic resolution at which performance is evaluated. We revisit these evaluation choices for ICU mortality prediction on MIMIC-IV, comparing predictive-utility and subgroup-error metrics across several fairness int...
A. Al Noman, Fahmid Al Rifat, Tahrima Hashem et al.· 0 citations
ICU mortality models can achieve strong discrimination, yet a risk score alone provides limited context for patient-level interpretation. We developed a multidimensional prediction-context framework that complements a calibrated mortality estimate with model behavior, data availability, recent physiology, and model att...
S. Gupta, A. Das, M. S. Anto et al.· medRxiv· 0 citations
A rigorous empirical framework is presented for comparing three uncertainty quantification approaches on two clinical prediction tasks, in-hospital mortality and 30-day readmission, using 74,829 ICU admissions from the MIMIC-IV database to support a more demanding evaluation standard for UQ in clinical machine learning...
Isaac Tosin Adisa, Francis Mawutor Amuyao, Ezekiel Olaoluwa Joaquim· International journal of re...· 0 citations
The Routing Disparity metric and multi-level debiasing framework introduced here generalize to MoE systems operating on demographically heterogeneous populations, providing both an audit tool and an architectural intervention for a bias mechanism that existing fairness methods leave unaddressed.
Xiaoyang Wang, Christopher C. Yang· IEEE journal of biomedical a...· 0 citations
It is established that population-level validation alone is insufficient for equity assessment of digital health AI, motivating subgroup-disaggregated reporting as a default standard, and subgroup-disaggregated reporting as a default standard for personalized configurations.
Junjie Luo, Xuzhe Zhi, Rui Han et al.· 0 citations
A fairness-aware machine learning framework using counterfactual adjustment to account for historical inequities embedded in clinical data reduced treatment disparities by 64.8% without compromising predictive accuracy (AUC = 0.89).
J. Kabangu, V. Lopera, Teresia M. Perkins et al.· Journal of Racial and Ethnic...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.