Aug 2026· Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2· 0 citations· 83 references
TL;DR
This paper introduces universal fairness, a clinically grounded definition that reframes fairness as maximizing subgroup-aware diagnostic performance under attribute-conditioned health disparities, and proposes MAPPE, a training-free minimax prompt optimization framework that theoretically promotes universal fairness.
Abstract
Large Language Models (LLMs) have shown strong potential in medical applications such as question answering and clinical prediction. % Despite their growing adoption, fairness in LLMs for medicine remains underexplored, largely due to the mismatch between conventional fairness constraints and the clinically meaningful role of sensitive attributes. Existing approaches often enforce attribute-invariant constraints, leading to substantial performance degradation that is unacceptable in high-stakes healthcare settings. Moreover, fairness evaluation for medical LLMs is hindered by the lack of dedicated benchmarks. In this paper, we first rethink fairness in LLMs for medicine from a clinically grounded, utility-based perspective. Inspired by principles of health equity in medicine, we introduce universal fairness, a clinically grounded definition that reframes fairness as maximizing subgroup-aware diagnostic performance under attribute-conditioned health disparities. To achieve this objective in practice, we propose MAPPE, a training-free minimax prompt optimization framework. % MAPPE theoretically promotes universal fairness, while directly applicable to both closed-source and open-source LLMs. To systematically evaluate fairness, we construct FairMed, the first attribute-annotated benchmark for medical LLMs covering medical question answering and clinical prediction. % Experiments on both closed-source and open-source LLMs reveal demographic disparities, while MAPPE consistently improves worst-group and overall performance, outperforming existing fairness-oriented and prompt-based methods. The dataset and code are available at https://github.com/xiye7lai/FairMed.
The Routing Disparity metric and multi-level debiasing framework introduced here generalize to MoE systems operating on demographically heterogeneous populations, providing both an audit tool and an architectural intervention for a bias mechanism that existing fairness methods leave unaddressed.
Xiaoyang Wang, Christopher C. Yang· IEEE journal of biomedical a...· 0 citations
It is argued that clinical utility, performance-based metrics, calibration, and statistical parity are the most relevant group-based metrics for medical applications and that different metrics might be applicable depending on the intended use and ethical framework.
S. L. van der Meijden, Yuqing Wang, Madelena Y. Ng et al.· The Lancet Digital Health· 1 citation
This work introduces and evaluates a privacy-preserving knowledge distillation framework for LLM-based clinical modeling, using multimorbidity scoring as a healthcare task, and establishes a trustworthy, privacy-compliant pathway for large-scale healthcare applications of LLMs.
R. Awasthi, Yi-He Yang, Meng-Xuan Li et al.· medRxiv· 0 citations
This paper introduces MaternaAI, a fairness-aware and explainable learning framework designed to enhance maternal healthcare predictions in Kerala, India, and proposes Adaptive Equity Score Optimization (AESO), a novel optimization algorithm that dynamically integrates fairness constraints into model training.
Large language models are increasingly used in contexts where their outputs can affect people directly, including hiring, admissions, and lending. This growing role makes it important to consider not only how well these models perform, but also whether their behavior is fair. Although many fairness metrics, bias benchm...
Samah Alhazmi· International Journal of Adv...· 0 citations
Results indicate that global feature importance, used as an active search signal rather than a post-hoc diagnostic, improves both the effectiveness and the efficiency of individual fairness testing.
H. Mamman, Abdullateef Oluwagbemiga Balogun, Mustapha Maidawa et al.· Journal of King Saud Univers...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.