Skip to content
Review Open access

Methods of Evaluating Large Language Model–Based Health Care Applications Used by Nonprofessionals: Protocol for a Scoping Review

Feb 2026 · JMIR Research Protocols · Vol 15, pp. e93509 · 0 citations · 38 references
Medicine

TL;DR

A scoping review that maps approaches for evaluation of LLM-based applications used for health purposes by nonprofessionals and will provide directional guidance for further research and development in the field of quality assurance for LLM-based applications used by nonprofessionals is outlined.

Abstract

Background Large language models (LLMs) are increasingly used in health care by nonprofessionals (ie, individuals without formal training in health-related professions). These applications must be evaluated in an appropriate manner to prevent misinformation and harmful decisions. To date, guidance to evaluate LLM-based applications for nonprofessional users remains limited and fragmented, leaving researchers and developers without a scientifically grounded set of quality dimensions, metrics, and measurement tools to guide them. Objective This protocol outlines a scoping review that maps approaches for evaluation of LLM-based applications used for health purposes by nonprofessionals. It identifies current methods and maps them thematically by assigning them to evaluation dimensions, metrics, and measurement instruments. The review will provide a comprehensive overview of evaluation methods currently in use. Methods The study follows the Joana Briggs Institute approach for conducting scoping reviews and reports. The protocol is reported in accordance with the PRISMA-P (Preferred Reporting Items for Systematic Reviews and Meta-Analyses Protocols) guidelines, and the scoping review will be reported in accordance with the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews) guidelines. The inclusion criteria comprise studies that evaluate LLM-based applications that are used in the context of health care by nonprofessionals. The search was conducted in PubMed, CINAHL, PsycInfo, and IEEE Xplore. Results since 2021 were considered. Data will be summarized and interpreted qualitatively. Publication screening was conducted by 2 independent reviewers in a blinded manner, with discrepancies settled through discussion. Data extraction and charting will be performed by 1 reviewer. To ensure quality, a random 10% sample of the publications will be independently charted by a second reviewer. Disagreements in the double-extracted subset will be resolved through discussion. Results As of July 2026, a steering committee of 6 researchers has been chosen for the conduct of the review. An initial search resulted in 8538 records after removing duplicates. After screening of these 8538 publications, 17.8% (1524/8538) were eligible for retrieval, of which 88.3% (1345/1524) were retrieved. Full-text screening (completed by 1 reviewer) excluded publications due to nonmatching populations (155/1345, 11.5%), concepts (246/1345, 18.3%), and contexts (24/1345, 1.8%), as well as secondary work (14/1345, 1%), leaving 67.4% (906/1345) of these publications for data extraction. We plan to perform final full-text screening, data extraction, coding, and synthesis of results in the fourth quarter of 2026. Conclusions The scoping review aims to identify and map current evaluation methods for LLM-based applications used in health care by nonprofessionals. It will provide a systematic overview of the current state of research and insights into quality dimensions, metrics, and measurement instruments. The findings will provide directional guidance for further research and development in the field of quality assurance for LLM-based applications used by nonprofessionals. International Registered Report Identifier (IRRID) DERR1-10.2196/93509

Read PDF

Similar papers

Review Aug 2026

How large language models can be used for teamwork and communication in healthcare settings: A scoping review.

LLMs hold substantial potential to enhance healthcare teamwork by supporting clinical decisions, streamlining administrative workflows, and improving patient communication, however, ethical, legal, and accountability concerns remain.

Ilse Super, Olya Rezaeian, Onur Asan · 0 citations
Review Jul 2026

Do Large Language Models Support Nursing Care Planning? A Scoping Review of Applications, Evaluation Approaches and Outcomes.

AIM To examine the overall performance of large language models (LLMs) in generating nursing care plans, clarify their role in nursing practice and identify directions for future research. DESIGN This study conducted a scoping review in accordance with Arksey and O'Malley's methodological framework. METHODS Five electronic databases were systematically searched: Web of Science Core Collection, PubMed, Scopus, CINAHL and IEEE Xplore. The search was limited to studies published between 1 June 2018 and 5 April 2026. RESULTS Fifteen studies were included. Existing studies primarily used nonreal patient cases and evaluated the textual quality of model-generated nursing care plans across a range of specialties. None examined LLM use within real-world clinical nursing workflows. Evaluation criteria mainly focused on accuracy, information quality and reliability, and readability. The strengths of LLMs in nursing care planning were concentrated in text organization, standardized terminology matching, and the initial drafting of nursing goals and interventions. However, important challenges remain, including privacy, hallucination, and bias. CONCLUSIONS LLMs may serve as assistive tools for generating initial drafts of nursing care plans, but they cannot yet replace nurses' clinical judgement. Future research should further refine evaluation frameworks and examine the impact of LLM-generated nursing care plans within real-world nursing workflows. Nurse-led human-AI collaboration should be emphasized to support the responsible translation of LLM-assisted nursing care planning into practice. IMPACT This scoping review highlights that, at present, LLMs can only serve as assistive tools in the development of nursing care plans, while nurses remain the primary decision-makers. It also underscores the need to enhance nurses' AI literacy to strengthen human-AI collaboration and facilitate the integration of LLMs as valuable supportive tools in nursing practice. PATIENT OR PUBLIC CONTRIBUTION No patient or public contribution.

Jianwen Zeng, Xule Zhu, Shiying Shen et al. · 0 citations
Review Open access Aug 2026

The Reliability of Human Evaluation of Large Language Models in Health Care Settings: Scoping Review.

It is suggested that the reliability of health care LLM reliability is difficult to evaluate adequately using a single universal standard, and future evaluations of health care LLM reliability need to be guided by standardized evaluation frameworks that reflect domain-specific contexts.

Euijun Yang, S. Ko, Hyekyung Woo · 0 citations
Review Open access Aug 2026

Safety-Oriented Evaluation of Large Language Models in Health Care: Guideline-Informed Systematic Review

Current evaluations emphasize accuracy while underreporting privacy, security, robustness, explainability, and verifiability, underscoring the need for comprehensive, guideline-based, multidomain evaluation before deployment in high-stakes clinical settings.

Hikaru Matsuoka, Takayuki Takahashi, Takayuki Semitsu et al. · 0 citations
Review Open access Jul 2026

An Acceptance Criteria Framework for Determining the Implementation Fit of Custom Large Language Models in Public Health Interventions

Abstract Large language models (LLMs) are increasingly embedded in clinical and population health workflows, including conversational agents such as health chatbots. As chatbots evolve from rule-based approaches to hybrid and LLM-enabled designs, risks and concerns about deployment readiness shift. Unlike rule-based chatbots, LLM outputs can be unpredictable, error-prone, and difficult to validate with traditional evaluation methods. Public health teams integrating customized LLMs into interventions face practical and ethical challenges related to performance variability, uncertainties about model behaviors, and inequitable performance across languages. Although existing frameworks address domains such as safety, ethics, effectiveness, engagement, and implementation, they often assume or imply—rather than operationalize—an explicit benchmark for deployment and implementation decisions. We propose an acceptance criteria framework (ACF) to determine implementation fit, defined as meeting prespecified minimum performance standards and demonstrating nonproblematic behavior under anticipated use. The ACF uses project-relevant and off-topic prompts, structured expert review, and prespecified thresholds to produce a documented decision record that can be iteratively rerun after model revisions. We demonstrate the framework through a case application in a tobacco cessation text messaging intervention, illustrating how the ACF can guide deployment decisions.

Andy J. King, Anthony Banks, L. Hernández et al. · 0 citations
Review Open access Aug 2026

Large Language Models for Mental Health Prediction: Scoping Review of Bias and Clinical Utility Documentation in 2019-2024

Abstract Background A growing body of literature leverages large language models (LLMs) to make mental health predictions. However, these models are prone to bias, and studies to validate their clinical utility are lacking. Objective This scoping review aims to uncover bias and clinical utility limitations stemming from the methodological design of LLM-based mental health predictive systems. In addition, it intends to document the level of self-reflection about bias and clinical challenges reported by authors in their own work. Methods This work follows the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews) guidelines and was registered online. Eligible studies were original research articles in English published between 2019 and 2024, using LLMs to detect mental health conditions in nonsynthetic textual data. The search was conducted in 5 scientific databases (PubMed, Web of Science, IEEE Xplore, ACM Digital Library, and ACL Anthology) with queries associating keywords related to “Mental Health,” “Large Language Models,” and “Prediction.” We extracted both methodological information about the included studies and authors’ statements relevant to issues of bias and clinical utility. This extraction was based on a framework screening the entire pipeline of development of LLMs with applications in mental health: research design and selection, data collection, outcome definition, model development, and postdeployment considerations. Statistical description of the retrieved entities, as well as thematic coding, was performed for analysis. Results A total of 2472 articles were identified, of which 263 (10.6%) were assessed for eligibility, and 201 (8.1%) were included in the review. Included studies were mostly recent, indicating a growing interest in the use of LLMs for mental health predictions. Our analysis revealed that a majority of studies share similar methodological choices along their development pipeline: most of them focus on depressive disorders identified via processing user texts on social media, mainly with the use of nonspecialist LLMs derived from BERT (Bidirectional Encoder Representations from Transformers). Following previous works on these matters, we highlighted how these choices may hinder the clinical relevance and fairness of the envisioned systems. Similarly, we found that 164 (81.6%) studies mention themes related to bias and clinical utility; however, most of the discussion revolves around data-centered issues. Only 41 (20.4%) articles mention themes associated with at least 3 out of 5 pipeline steps, suggesting a limited appropriation of the notions of bias and clinical utility in such a sensitive context as mental health analysis. Conclusions Bias and clinical utility are lightly covered in the field of LLM-based mental health prediction research as of 2019‐2024. In-depth approaches involving interdisciplinary teams of clinicians and natural language processing specialists are needed to ensure technical soundness, clinical relevance, and fair outcomes for potential users.

Clémentine Bleuze, Karen Fort, Vincent P. Martin et al. · 0 citations