Skip to content
Review Open access

Evaluating large language models as clinical decision support tools in primary healthcare settings: Protocol for a multi-country comparative validation study on expert-adjudicated hypothetical vignettes (hypMOOVE-PHC)

Sep 2026 · medRxiv · 0 citations
Medicine

TL;DR

The hypMOOVE-PHC study is the hypothetical vignette phase of the Massive Open Online Validation and Evaluation (MOOVE) initiative, implemented in Kenya, Malawi, and Tanzania, and aims to validate a pool of LLMs through clinical review of expert-generated vignettes through fully crossed repeated-measures comparative evaluation study.

Abstract

Introduction: Large language models (LLMs) have the potential to strengthen clinical decision-making in low-resource primary healthcare (PHC) settings. However, most LLMs are developed and benchmarked in high-resource settings and evidence on their safety and contextual appropriateness in Sub-Saharan Africa remains limited. The hypMOOVE-PHC study is the hypothetical vignette phase of the Massive Open Online Validation and Evaluation (MOOVE) initiative, implemented in Kenya, Malawi, and Tanzania. It aims to validate a pool of LLMs through clinical review of expert-generated vignettes. Methods and analysis: This is a fully crossed repeated-measures comparative evaluation study. In each country, experienced clinicians develop 200-250 hypothetical clinical vignettes reflecting realistic patient presentations and independently produce a human benchmark care plan for each. Vignettes are used to prompt a selection of six open-source and proprietary LLMs selected based on code availability, local hostability, and model size. During in-person workshops (valiDATAthons), independent clinical experts rate LLM- and human-generated responses in source-attribution masked side-by-side comparisons across five dimensions (clinical soundness, safety, contextual fit, clarity & completeness, and appropriate confidence). The primary endpoints are each LLM's overall performance profile and non-inferior safety profile, as compared to the human benchmark. At minimum, 358 evaluations per LLM (or 1,253 paired evaluations in total) are required per country. Ethics and dissemination: The study is approved by the EPFL Human Ethics Research Committee in Switzerland, Harvard T.H. Chan School of Public Health in the USA, KNH-UoN Ethics and Research Committee in Kenya, MUBAS Research Ethics Committee in Malawi, and MUHAS Research and Ethics Committee and National Institute for Medical Research in Tanzania. Findings will be reported according to the TRIPOD-LLM framework and shared with national ministries of health, disseminated at conferences and in peer-reviewed journals, and de-identified benchmark data will be released under FAIR principles.

Read PDF

Similar papers

#large language models Review Open access Sep 2026

Quality and safety of large language model–generated medication review outputs in geriatric pharmacotherapy: a two-stage comparative vignette-based benchmark evaluation

These findings support supervised use of large language models and evaluation approaches that assess reasoning and prioritisation as well as target detection and should not be interpreted as evidence that one model is clinically superior in real-world practice.

Kubra Cingar Alpay, D. Ozata, T. Gedik et al. · 0 citations
Open access Aug 2026

Benchmarking large language models for HIV medical decision support

HIVMedQA is developed, a clinician-curated benchmark of HIV-related open-ended medical question-answer pairs spanning basic knowledge, clinical reasoning, complex patient vignettes, and bias-modified scenarios that provides a structured benchmark for evaluating LLMs in HIV clinical decision support.

Gonzalo Cardenal-Antolin, J. Fellay, Bashkim Jaha et al. · 1 citation
Review Open access Aug 2026

AI safety evaluation in an underrepresented population: real-world performance of clinical decision support and frontier language models on Medicaid patient messaging triage

No evaluated tool or combination was sufficiently accurate to enable physician-unassisted triage in this setting of patient-initiated text messages in a multi-state Medicaid population.

S. Basu, Sadiq Y. Patel, Parth Sheth et al. · 0 citations
Review Open access Aug 2026

Physician-revised AI-generated drafts are associated with higher ratings of written explanations in end-of-life care in the intensive care unit: a scenario-based single-center cross-sectional study

In end-of-life care (EOL) in the intensive care unit (ICU), intensivists are expected to provide medically appropriate and empathetic communication to support shared decision-making with patients and their families. Large language models (LLMs) have shown potential to generate medical responses that are perceived as in...

Atsushi Kokita, Junpei Haruna, Yuya Goto et al. · 0 citations
Review Open access Aug 2026

Benchmarking publicly accessible large language models for English-language patient-facing acute pancreatitis information: a cross-sectional study of quality, transparency, and readability

Background Patients increasingly rely on large language models (LLMs) for health information, yet their suitability for decision-critical conditions such as acute pancreatitis remains unclear. Given that acute pancreatitis requires timely symptom recognition, severity assessment, treatment decision-making, recurrence p...

Biao Jiang, Hong-Xin Sun, Linlin Chen · 0 citations
#explainable ai Review Open access Sep 2026

Large language models in healthcare: applications, evaluation frameworks, and governance pathways — a scoping review and multidimensional framework

Background Large language models (LLMs) are increasingly evaluated for healthcare applications spanning clinical documentation, decision support, patient communication, research assistance, and operational workflows. Despite rapid adoption interest, evidence remains heterogeneous and standardised approaches for evaluat...

J. C. Ferreira, Isabel Rosa · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.