Recommendations are provided that focus on improving transparency in reporting, advocating for the broader adoption of multi-level fairness techniques, and ensuring that health equity is explicitly prioritized in future research efforts.
Abstract
The increasing integration of machine learning in healthcare has highlighted critical challenges related to fairness, transparency, and health equity. Specifically, the use of multi-level fairness techniques, which combine multiple bias mitigation steps or techniques, show promise for reducing biases across different patient demographics, yet this approach remains underexplored in terms of its health equity outcomes. In this paper, we assess the current landscape of multi-level fairness in health informatics by focusing on its impact on equitable healthcare outcomes and evaluating how transparency and reporting standards contribute to these advancements. Through an examination of the existing literature, we identify key gaps in both the implementation of multi-level fairness techniques and the consistent reporting of health equity impacts. Furthermore, we analyze the role of reporting standards, including MINIMAR and TRIPOD, in improving model transparency and ensuring that machine learning models in healthcare address health disparities. These standards offer valuable benchmarks for reporting on ML models, yet we identify key opportunities for enhancing how these reports capture fairness and equity outcomes. The paper concludes by providing recommendations that focus on improving transparency in reporting, advocating for the broader adoption of multi-level fairness techniques, and ensuring that health equity is explicitly prioritized in future research efforts.
Maternal healthcare prediction systems often suffer from algorithmic biases due to socio-economic disparities and imbalanced datasets, limiting their effectiveness for equitable healthcare policymaking. This paper introduces MaternaAI, a fairness-aware and explainable learning framework designed to enhance maternal healthcare predictions in Kerala, India. The framework focuses on three critical health indicators:(1) Tetanus Toxoid (TT) booster uptake,(2) immunization coverage rates, and (3) the percentage of pregnant women completing four or more Antenatal Care (ANC) visits. To address fairness, we propose Adaptive Equity Score Optimization (AESO), a novel optimization algorithm that dynamically integrates fairness constraints into model training. AESO is model-agnostic and adapts group equity weights in response to real-time disparities. We integrate SHAP, LIME, and feature permutation techniques for explainability, enabling transparent global and local interpretation. Empirical results demonstrate that MaternaAI significantly improves fairness metrics and model accuracy across diverse machine learning and deep learning models, offering interpretable and equitable decision support for public health stakeholders.
Artificial intelligence (AI) is increasingly embedded in health systems globally and has the potential to improve efficiency, diagnostic accuracy, and decision support. However, its benefits remain unevenly distributed, particularly in low- and middle-income countries (LMICs). Models developed using datasets from specific populations may perform poorly in other settings, reinforcing structural inequities rather than correcting them. This viewpoint proposes a composite framework, the AI in Healthcare Equity Index (AIHEI), to support measurable assessment of equity in health AI systems. The AIHEI is designed to assess equity across five domains: data representation, algorithmic fairness, transparency and explainability, governance and oversight, and community impact and benefit sharing. By generating a standardised score, the index could enable comparisons across technologies, incentivise improvement, and support regulation, procurement, publication, and funding decisions. Pilots across diverse health domains and geographic settings are needed to assess feasibility, refine domain weighting, and evaluate reliability, reproducibility, and validity. Important challenges include contextual definitions of fairness, data sovereignty, post-deployment monitoring, and the risk of metric gaming. Quantifying equity in health AI is essential to ensure that AI does not create, widen, or exacerbate existing disparities by neglecting underserved populations. A common, objective measure of AI-related health equity can help move the field from ethical aspiration toward measurable accountability, monitoring, and enforcement.
Basile Njei, U. S. Kanmounye, L. Bain et al.· International Journal for Eq...· 0 citations
Large Language Models (LLMs) have shown strong potential in medical applications such as question answering and clinical prediction. % Despite their growing adoption, fairness in LLMs for medicine remains underexplored, largely due to the mismatch between conventional fairness constraints and the clinically meaningful role of sensitive attributes. Existing approaches often enforce attribute-invariant constraints, leading to substantial performance degradation that is unacceptable in high-stakes healthcare settings. Moreover, fairness evaluation for medical LLMs is hindered by the lack of dedicated benchmarks. In this paper, we first rethink fairness in LLMs for medicine from a clinically grounded, utility-based perspective. Inspired by principles of health equity in medicine, we introduce universal fairness, a clinically grounded definition that reframes fairness as maximizing subgroup-aware diagnostic performance under attribute-conditioned health disparities. To achieve this objective in practice, we propose MAPPE, a training-free minimax prompt optimization framework. % MAPPE theoretically promotes universal fairness, while directly applicable to both closed-source and open-source LLMs. To systematically evaluate fairness, we construct FairMed, the first attribute-annotated benchmark for medical LLMs covering medical question answering and clinical prediction. % Experiments on both closed-source and open-source LLMs reveal demographic disparities, while MAPPE consistently improves worst-group and overall performance, outperforming existing fairness-oriented and prompt-based methods. The dataset and code are available at https://github.com/xiye7lai/FairMed.
Jiaming Zhang, Yuyuan Li, Xiaohua Feng et al.· Proceedings of the 32nd ACM...· 0 citations
Federated learning (FL) enables multi-institutional collaboration in medical imaging while preserving patient privacy, yet its fairness landscape remains fragmented: existing methods predominantly address either
collaboration fairness
(equitable performance across institutions) or
group fairness
(equitable outcomes across demographic subgroups), but rarely both. In this systematic review, we adopt
dual fairness
—the joint satisfaction of both dimensions—as the analytical lens for organizing and critically evaluating this landscape. Following the PRISMA 2020 guidelines, we analyze 132 publications and classify fairness-aware FL methods through a three-dimensional taxonomy: client-side, server-side, and communication-based approaches. Among the 20 fairness-aware or fairness-adapted FL methods catalogued, only three partially address both dimensions, and none provides provable joint guarantees under clinically realistic conditions. Our critical analysis identifies three fundamental challenges: the Local–Global Pareto Frontier Conflict, in which collaboration and group fairness gradients in the accuracy space can exceed
$$150^{\circ }$$
under sufficiently asymmetric demographic imbalance (e.g.,
$$\pi _1 \ge 0.8$$
); the Privacy–Fairness Compounding Effect, through which differential privacy mechanisms disproportionately suppress minority gradient signals; and the risk of pseudo-fairness, whereby equipment–demographic confounding masks genuine algorithmic discrimination. We further outline a seven-direction research roadmap. To the best of our knowledge, this constitutes the first systematic review to formally analyze the gradient-level conflict between collaboration fairness and group fairness in federated medical imaging, while also providing a structured causal analysis of equipment–demographic confounding, offering both a critical synthesis and actionable directions toward equitable AI-assisted healthcare.
INTRODUCTION
While health-care quality improvement (QI) initiatives have improved health outcomes, these benefits are not accessible to all patients. Currently, there is no clear roadmap for implementing QI to address health inequities. Therefore, we studied the experiences of collaborative quality improvement (CQI) leaders who leverage QI to address health inequities.
METHODS
We conducted semi-structured group and individual interviews with 25 leaders from 12 CQI coordinating centers in Michigan between May and August 2025. Participants were asked about their current QI initiatives that seek to explore or mitigate the impact of nonmedical drivers of health outcome variability. Interviews were audio-recorded, transcribed verbatim, and deidentified. Data were analyzed using a combination of deductive and inductive content analysis, followed by a secondary matrix analysis.
RESULTS
Participants identified challenges and opportunities in conducting QI to evaluate gaps and ensure access to quality care across three major themes. First, data points and variables that assess nonmedical drivers of care were often missing or lacked standardization. Second, CQIs and hospitals had an inconsistent understanding of health equity. Lastly, patients and hospitals had variable access to resources that could help to address gaps in access and improve quality.
CONCLUSIONS
In this study, CQI leaders identified common challenges that affect their ability to ensure that their QI efforts are accessible and effective for all patients. By collaboratively learning from these challenges and sharing the paths they identified to move forward, there is greater potential to ensure that QI can provide access to high-quality care for all individuals.
Jamila K. Picart, Erin E. Isenberg, Sarah E. Bradley et al.· Journal of Surgical Research· 0 citations
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.
MIT News · Artificial Intelligence· news.mit.eduAug 24, 2026