Aug 2026· Journal of Medical Internet Research· Vol 28, pp.
e98184
· 0 citations· 121 references
Medicine
TL;DR
It is suggested that the reliability of health care LLM reliability is difficult to evaluate adequately using a single universal standard, and future evaluations of health care LLM reliability need to be guided by standardized evaluation frameworks that reflect domain-specific contexts.
Abstract
Background
Integration of large language models (LLMs) into health care has accelerated rapidly, yet reliability concerns pose potential risks to patient safety. Although human evaluation has been widely used as an important approach for assessing LLM reliability, a systematic understanding of how such evaluations have been operationalized across studies remains limited.
Objective
This study aimed to characterize the current landscape of human evaluation frameworks for LLM reliability in health care and to identify similarities and differences between the clinical and public health domains.
Methods
In line with the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses Extension for Scoping Reviews) guidelines, PubMed, Web of Science, the Cochrane Library, CINAHL, and Google Scholar were searched for studies published from January 2016 to July 2025. Eligible studies were English-language original research conducted in health care settings that assessed the reliability of LLM-generated responses through human evaluation. Key exclusion criteria were studies without human evaluation and studies focused primarily on LLM model selection, performance optimization, or technical development. Extracted data were analyzed across 3 dimensions: what was evaluated, who evaluated, and how evaluation was conducted. Reported methodological limitations were also categorized and compared between the clinical and public health domains.
Results
Of the 4347 records identified, 71 studies were included in the final analysis (clinical, n=26; public health, n=45). Six reliability indicators were used: accuracy, relevance, completeness, clarity, safety, and consistency. The clinical domain more frequently assessed guideline concordance, internal consistency, and structural coherence, whereas the public health domain more frequently assessed understandability, harm potential, and repeat response consistency. Single-specialty clinicians were the most common evaluators in both domains, although mixed evaluator panels were observed only in the public health domain. Evaluator panels generally consisted of 5 or fewer members. Five-point Likert scales and researcher-defined rubrics were commonly used evaluation approaches in both domains. Key methodological limitations included evaluator subjectivity, nonstandardized indicators, and limited evaluation scope and settings.
Conclusions
To our knowledge, this is the first review to systematically examine how human evaluations of LLM reliability have been conducted across health care. The focus of reliability evaluation differed across domains, with clinical evaluations giving relatively greater attention to clinical validity and logical rigor, whereas public health evaluations gave relatively greater attention to understandability, practical use, and safe use. These differences suggest that the reliability of health care LLMs is difficult to evaluate adequately using a single universal standard. In addition, the methodological limitations identified in this review indicate that current human evaluation approaches are insufficiently standardized. Therefore, future evaluations of health care LLM reliability need to be guided by standardized evaluation frameworks that reflect domain-specific contexts and encompass indicator definitions, judgment criteria, evaluator guidance, and evaluation procedures.
Abstract Background A growing body of literature leverages large language models (LLMs) to make mental health predictions. However, these models are prone to bias, and studies to validate their clinical utility are lacking. Objective This scoping review aims to uncover bias and clinical utility limitations stemming from the methodological design of LLM-based mental health predictive systems. In addition, it intends to document the level of self-reflection about bias and clinical challenges reported by authors in their own work. Methods This work follows the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews) guidelines and was registered online. Eligible studies were original research articles in English published between 2019 and 2024, using LLMs to detect mental health conditions in nonsynthetic textual data. The search was conducted in 5 scientific databases (PubMed, Web of Science, IEEE Xplore, ACM Digital Library, and ACL Anthology) with queries associating keywords related to “Mental Health,” “Large Language Models,” and “Prediction.” We extracted both methodological information about the included studies and authors’ statements relevant to issues of bias and clinical utility. This extraction was based on a framework screening the entire pipeline of development of LLMs with applications in mental health: research design and selection, data collection, outcome definition, model development, and postdeployment considerations. Statistical description of the retrieved entities, as well as thematic coding, was performed for analysis. Results A total of 2472 articles were identified, of which 263 (10.6%) were assessed for eligibility, and 201 (8.1%) were included in the review. Included studies were mostly recent, indicating a growing interest in the use of LLMs for mental health predictions. Our analysis revealed that a majority of studies share similar methodological choices along their development pipeline: most of them focus on depressive disorders identified via processing user texts on social media, mainly with the use of nonspecialist LLMs derived from BERT (Bidirectional Encoder Representations from Transformers). Following previous works on these matters, we highlighted how these choices may hinder the clinical relevance and fairness of the envisioned systems. Similarly, we found that 164 (81.6%) studies mention themes related to bias and clinical utility; however, most of the discussion revolves around data-centered issues. Only 41 (20.4%) articles mention themes associated with at least 3 out of 5 pipeline steps, suggesting a limited appropriation of the notions of bias and clinical utility in such a sensitive context as mental health analysis. Conclusions Bias and clinical utility are lightly covered in the field of LLM-based mental health prediction research as of 2019‐2024. In-depth approaches involving interdisciplinary teams of clinicians and natural language processing specialists are needed to ensure technical soundness, clinical relevance, and fair outcomes for potential users.
Clémentine Bleuze, Karen Fort, Vincent P. Martin et al.· JMIR AI· 0 citations
Current evaluations emphasize accuracy while underreporting privacy, security, robustness, explainability, and verifiability, underscoring the need for comprehensive, guideline-based, multidomain evaluation before deployment in high-stakes clinical settings.
Hikaru Matsuoka, Takayuki Takahashi, Takayuki Semitsu et al.· Online Journal of Public Hea...· 0 citations
LLMs hold substantial potential to enhance healthcare teamwork by supporting clinical decisions, streamlining administrative workflows, and improving patient communication, however, ethical, legal, and accountability concerns remain.
Ilse Super, Olya Rezaeian, Onur Asan· International Journal of Med...· 0 citations
Reporting transparency in radiology and medical imaging LLM studies published in 2025 was inconsistent across reporting items and journals, with substantial deficiencies in some reproducibility-critical elements.
I. Mese, Saime Turgut Gunes, Ozge Coskun et al.· Korean Journal of Radiology· 0 citations
A scoping review that maps approaches for evaluation of LLM-based applications used for health purposes by nonprofessionals and will provide directional guidance for further research and development in the field of quality assurance for LLM-based applications used by nonprofessionals is outlined.
Maren Keuchel, Pinar Bisgin, Tom Strube et al.· JMIR Research Protocols· 0 citations