Quality and safety of large language model–generated medication review outputs in geriatric pharmacotherapy: a two-stage comparative vignette-based benchmark evaluation
These findings support supervised use of large language models and evaluation approaches that assess reasoning and prioritisation as well as target detection and should not be interpreted as evidence that one model is clinically superior in real-world practice.
Abstract
Polypharmacy and multimorbidity make medication review in older adults a high-stakes clinical task. Large language models (LLMs) may support medication review, but whether item-level concordance with explicit prescribing criteria reflects expert-rated reasoning quality and safety prioritisation is uncertain. We compared the quality and safety of outputs generated by three LLMs using a two-stage evaluation framework.
In this two-stage comparative vignette-based benchmark evaluation, GPT-5.2, Claude Sonnet 4.5, and Gemini 3 Pro responded to 20 standardised geriatric pharmacotherapy vignettes representing fictional older adults aged 72–88 years across four clinical domains. Each vignette contained three potentially inappropriate medications and one clinically relevant START-type omission anchored to the AGS Beers Criteria and STOPP/START version 3. Models received an identical master prompt under default end-user settings; memory features were disabled where available and separate sessions were used for each vignette. Two geriatricians, blinded to model identity, independently rated anonymised outputs using a 100-point rubric comprising medication review output quality (0–80) and critical safety-risk prioritisation (0–20). Stage 1 assessed answer-key concordance; exploratory Stage 2 characterised clinically contextualised error patterns. Comparisons used Friedman and Wilcoxon signed-rank tests, Cochran’s Q, and exact McNemar tests with Bonferroni correction.
Stage 1 concordance was uniformly high, with limited between-model discrimination. Expert-rated total scores differed across models (
p
< 0.001; Kendall’s W = 0.700): Gemini 3 Pro scored highest (97.93 ± 2.33; 95% CI 96.83–99.02), followed by Claude Sonnet 4.5 (93.90 ± 3.37) and GPT-5.2 (89.85 ± 5.23); all pairwise comparisons were significant, with large standardised effect sizes (
r
= 0.55–0.61). Safety scores showed the same ranking (
p
< 0.001; Kendall’s W = 0.861;
r
= 0.52–0.62). Reliability of averaged total ratings was high (ICC(A,2) = 0.935; 95% CI 0.897–0.957). In the exploratory Stage 2 error analysis, Gemini was less frequently flagged than GPT-5.2 for superficial reasoning, weak emphasis on life-threatening risk, and any flagged error, and less frequently than Claude Sonnet 4.5 for weak emphasis on life-threatening risk and any flagged error.
Within this vignette-based benchmark, high item-level concordance did not ensure high expert-rated output quality or safety prioritisation. Models also differed in safety prioritisation and in the pattern of flagged errors. These findings support supervised use of LLMs and evaluation approaches that assess reasoning and prioritisation as well as target detection. They should not be interpreted as evidence that one model is clinically superior in real-world practice.
Clinical trial not applicable.
The hypMOOVE-PHC study is the hypothetical vignette phase of the Massive Open Online Validation and Evaluation (MOOVE) initiative, implemented in Kenya, Malawi, and Tanzania, and aims to validate a pool of LLMs through clinical review of expert-generated vignettes through fully crossed repeated-measures comparative eva...
P. Macharia, C. Kachimanga, M. Mahende et al.· medRxiv· 0 citations
Background General-purpose large language models are increasingly used by patients and caregivers to obtain mental health information and guidance about when professional care is required. In late-life depression, broadly accurate information may nevertheless be unsafe when cognitive change, multimorbidity, frailty, po...
Wei Xiao, Huan Zhang, Xiao-Yi Chen et al.· Frontiers in Psychiatry· 0 citations
Exam-style accuracy does not establish whether large language models (LLMs) reason well over clinical records. We define clinical reasoning as integrating and updating evidence across time and sources to form, revise and justify a patient's problem representation and a defensible plan. This structured narrative review...
Large language models (LLMs) can produce fluent medication explanations, draft patient-facing materials, summarize records, and assist with information retrieval. Their linguistic competence has encouraged proposals for pharmaceutical-care use, yet fluency is not equivalent to clinical reliability. This integrative nar...
Júlia Costa Oliveira Ornelas· Nexus Science Review· 0 citations
Large Language Models (LLMs) are increasingly consulted by millions of patients seeking pharmaceutical information, yet their reliability and safety in providing medication-related advice remain inadequately evaluated. This study assessed the performance of four leading LLMs in responding to frequently asked questions...
Y. L. T. Bayala, I. A. Tinni, Thierry Boris Wend-Yam Yaméogo et al.· Discover Artificial Intellig...· 0 citations
We propose a rubric for assessing expressed clinical reasoning in model responses, drawing on three bodies of work: medical education assessment frameworks (ART, SCT, Key Feature Problems and OSCE); clinical LLM benchmarks (MedR-Bench, HealthBench, TIMER-Bench, DR. BENCH, PrIME-LLM and PatientSafeBench); and general LL...
Zhang Jiang, Zina Ibrahim, J. Teo· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduOct 8, 2026