Skip to content
Review Open access

Safety-Oriented Evaluation of Large Language Models in Health Care: Guideline-Informed Systematic Review

Aug 2026 · Online Journal of Public Health Informatics · Vol 18 · 0 citations · 28 references
Medicine

TL;DR

Current evaluations emphasize accuracy while underreporting privacy, security, robustness, explainability, and verifiability, underscoring the need for comprehensive, guideline-based, multidomain evaluation before deployment in high-stakes clinical settings.

Abstract

Abstract Background Large language models (LLMs) are rapidly emerging in health care, offering opportunities in decision support, education, and research, but raising critical concerns about safety, reliability, and ethics. Although several guidelines for trustworthy AI exist in business and technology, few systematic reviews have applied them to medical contexts. Objective This study aimed to conduct a systematic review of LLM research in health care, applying the AI Guidelines for Business as a framework across 11 domains, including safety, reliability, ethics, transparency, fairness, inclusiveness, privacy, security, robustness, data quality, and verifiability. Methods Following PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) 2020 guidelines (retrospectively registered on the Open Science Framework; DOI 10.17605/OSF.IO/P4KSB), the PubMed, Scopus, Web of Science, arXiv, and IEEE Xplore databases were searched on January 15, 2025. Records were screened in 2 stages by 3 reviewers (with records retained only upon unanimous agreement). A total of 247 studies were included, of which 211 (85.4%) contributed quantitative values. Eligible studies were classified across 11 trustworthy AI domains. Heterogeneous metrics were summarized within metric families; when multiple models were evaluated, the mean across models was used as the primary estimate, with best, median, and primary-model sensitivity analyses. The LLM-assisted categorization (GPT-5 mini) was validated by using an automated internal consistency check, and 95% CIs were estimated by using cluster bootstrap on study-level values. Results Of the 25,156 records, 247 (1.0%) studies were included, and of these, 211 (85.4%) contributed quantitative values. Evaluation concentrated on accuracy (143/247, 57.9%) and fairness and inclusiveness (93/247, 37.7%), followed by data quality (47/247, 19.0%) and prevention of misinformation (44/247, 17.8%). Normalized performance was moderate to high (accuracy mean 0.73, 95% CI 0.7-0.76; data quality: 0.64; prevention of misinformation: 0.82). Selecting the best-performing model inflated domain means by up to 0.05. Privacy protection (2/247, 0.8%) and security assurance (0/247, 0.0%) were almost entirely absent. Domain assignments were recoverable from objective metric types in 98.9% of values (Cohen κ=0.985). Conclusions Current evaluations emphasize accuracy while underreporting privacy, security, robustness, explainability, and verifiability. This finding reflects gaps in reporting rather than demonstrated poor performance, underscoring the need for comprehensive, guideline-based, multidomain evaluation before deployment in high-stakes clinical settings.

Read PDF

Similar papers

Review Open access Aug 2026

The Reliability of Human Evaluation of Large Language Models in Health Care Settings: Scoping Review.

It is suggested that the reliability of health care LLM reliability is difficult to evaluate adequately using a single universal standard, and future evaluations of health care LLM reliability need to be guided by standardized evaluation frameworks that reflect domain-specific contexts.

Euijun Yang, S. Ko, Hyekyung Woo · 0 citations
Review Open access Aug 2026

Large Language Models for Mental Health Prediction: Scoping Review of Bias and Clinical Utility Documentation in 2019-2024

Abstract Background A growing body of literature leverages large language models (LLMs) to make mental health predictions. However, these models are prone to bias, and studies to validate their clinical utility are lacking. Objective This scoping review aims to uncover bias and clinical utility limitations stemming from the methodological design of LLM-based mental health predictive systems. In addition, it intends to document the level of self-reflection about bias and clinical challenges reported by authors in their own work. Methods This work follows the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews) guidelines and was registered online. Eligible studies were original research articles in English published between 2019 and 2024, using LLMs to detect mental health conditions in nonsynthetic textual data. The search was conducted in 5 scientific databases (PubMed, Web of Science, IEEE Xplore, ACM Digital Library, and ACL Anthology) with queries associating keywords related to “Mental Health,” “Large Language Models,” and “Prediction.” We extracted both methodological information about the included studies and authors’ statements relevant to issues of bias and clinical utility. This extraction was based on a framework screening the entire pipeline of development of LLMs with applications in mental health: research design and selection, data collection, outcome definition, model development, and postdeployment considerations. Statistical description of the retrieved entities, as well as thematic coding, was performed for analysis. Results A total of 2472 articles were identified, of which 263 (10.6%) were assessed for eligibility, and 201 (8.1%) were included in the review. Included studies were mostly recent, indicating a growing interest in the use of LLMs for mental health predictions. Our analysis revealed that a majority of studies share similar methodological choices along their development pipeline: most of them focus on depressive disorders identified via processing user texts on social media, mainly with the use of nonspecialist LLMs derived from BERT (Bidirectional Encoder Representations from Transformers). Following previous works on these matters, we highlighted how these choices may hinder the clinical relevance and fairness of the envisioned systems. Similarly, we found that 164 (81.6%) studies mention themes related to bias and clinical utility; however, most of the discussion revolves around data-centered issues. Only 41 (20.4%) articles mention themes associated with at least 3 out of 5 pipeline steps, suggesting a limited appropriation of the notions of bias and clinical utility in such a sensitive context as mental health analysis. Conclusions Bias and clinical utility are lightly covered in the field of LLM-based mental health prediction research as of 2019‐2024. In-depth approaches involving interdisciplinary teams of clinicians and natural language processing specialists are needed to ensure technical soundness, clinical relevance, and fair outcomes for potential users.

Clémentine Bleuze, Karen Fort, Vincent P. Martin et al. · 0 citations
Review Open access Aug 2026

Privacy, security, and reliability risks of artificial intelligence in healthcare: a systematic review of empirical evidence.

This systematic review provides empirical evidence suggesting that contemporary AI systems in healthcare introduce privacy and security risks that may challenge traditional assumptions about data protection, and underscores the need for privacy- and security-by-design approaches and governance frameworks that address risks across the AI lifecycle.

Elvin Khanjahani, Shabnam Iezadi, Savannah Marshall et al. · 0 citations
Review Open access Feb 2026

Methods of Evaluating Large Language Model–Based Health Care Applications Used by Nonprofessionals: Protocol for a Scoping Review

A scoping review that maps approaches for evaluation of LLM-based applications used for health purposes by nonprofessionals and will provide directional guidance for further research and development in the field of quality assurance for LLM-based applications used by nonprofessionals is outlined.

Maren Keuchel, Pinar Bisgin, Tom Strube et al. · 0 citations
Review Aug 2026

Ethical and Regulatory Issues in Governing AI-Enabled Software as a Medical Device: A Scoping Review.

PURPOSE Artificial intelligence (AI)-enabled software as a medical device (SaMD) is increasingly used across clinical specialties, but its governance remains difficult because adaptive systems raise ongoing concerns about version control, subgroup performance reporting, and postmarket performance drift. This scoping review examined the ethical and regulatory issues surrounding AI-enabled SaMD across jurisdictions and across the product lifecycle, drawing together peer-reviewed literature, standards, and official guidance. METHODS A scoping review was conducted using the Joanna Briggs Institute framework and reported in accordance with PRISMA-ScR. Sources published between 2015 and 2025 were identified from PubMed, PubMed Central, Google Scholar, medRxiv/bioRxiv, and the Cochrane Library. Included records were charted by lifecycle stage, jurisdiction, and six prespecified themes: transparency, equity, privacy and security, accountability, lifecycle oversight and change control, and convergence versus fragmentation. Coding was refined iteratively during reviewer calibration. FINDINGS Of 21,672 records identified, 369 met the inclusion criteria. Most sources were published from 2019 onward and were concentrated in United States and European Union regulatory settings, with additional contributions from international bodies such as International Medical Device Regulators Forum, WHO, and ISO, as well as from emerging economies including China and India. Postmarket oversight emerged as the most strongly emphasized lifecycle stage, especially in relation to real-world monitoring, drift management, and prespecified change control. Common gaps included limited subgroup reporting, inconsistent expectations for drift thresholds and rollback criteria, and poor alignment between horizontal AI rules and device-specific regulatory frameworks. IMPLICATIONS The evidence base shows meaningful progress in the governance of AI-enabled SaMD, but implementation remains uneven across jurisdictions and lifecycle stages. Priority areas include standardized equity reporting, clearer minimum expectations for drift management, and more explicit integration between horizontal AI governance frameworks and SaMD-specific regulatory requirements. These findings support stronger accountability through improved reporting standards, clearer postmarket controls, and better alignment of quality-management processes across the AI-SaMD lifecycle.

Shaharyar Ahsan, Vivian Annastasia Chinyem Obi, D. Noor et al. · 0 citations