Skip to content
#large language models Review Open access

Quality and safety of large language model–generated medication review outputs in geriatric pharmacotherapy: a two-stage comparative vignette-based benchmark evaluation

Sep 2026 · BMC Geriatrics · 0 citations

TL;DR

These findings support supervised use of large language models and evaluation approaches that assess reasoning and prioritisation as well as target detection and should not be interpreted as evidence that one model is clinically superior in real-world practice.

Abstract

Polypharmacy and multimorbidity make medication review in older adults a high-stakes clinical task. Large language models (LLMs) may support medication review, but whether item-level concordance with explicit prescribing criteria reflects expert-rated reasoning quality and safety prioritisation is uncertain. We compared the quality and safety of outputs generated by three LLMs using a two-stage evaluation framework. In this two-stage comparative vignette-based benchmark evaluation, GPT-5.2, Claude Sonnet 4.5, and Gemini 3 Pro responded to 20 standardised geriatric pharmacotherapy vignettes representing fictional older adults aged 72–88 years across four clinical domains. Each vignette contained three potentially inappropriate medications and one clinically relevant START-type omission anchored to the AGS Beers Criteria and STOPP/START version 3. Models received an identical master prompt under default end-user settings; memory features were disabled where available and separate sessions were used for each vignette. Two geriatricians, blinded to model identity, independently rated anonymised outputs using a 100-point rubric comprising medication review output quality (0–80) and critical safety-risk prioritisation (0–20). Stage 1 assessed answer-key concordance; exploratory Stage 2 characterised clinically contextualised error patterns. Comparisons used Friedman and Wilcoxon signed-rank tests, Cochran’s Q, and exact McNemar tests with Bonferroni correction. Stage 1 concordance was uniformly high, with limited between-model discrimination. Expert-rated total scores differed across models ( p  < 0.001; Kendall’s W = 0.700): Gemini 3 Pro scored highest (97.93 ± 2.33; 95% CI 96.83–99.02), followed by Claude Sonnet 4.5 (93.90 ± 3.37) and GPT-5.2 (89.85 ± 5.23); all pairwise comparisons were significant, with large standardised effect sizes ( r  = 0.55–0.61). Safety scores showed the same ranking ( p  < 0.001; Kendall’s W = 0.861; r  = 0.52–0.62). Reliability of averaged total ratings was high (ICC(A,2) = 0.935; 95% CI 0.897–0.957). In the exploratory Stage 2 error analysis, Gemini was less frequently flagged than GPT-5.2 for superficial reasoning, weak emphasis on life-threatening risk, and any flagged error, and less frequently than Claude Sonnet 4.5 for weak emphasis on life-threatening risk and any flagged error. Within this vignette-based benchmark, high item-level concordance did not ensure high expert-rated output quality or safety prioritisation. Models also differed in safety prioritisation and in the pattern of flagged errors. These findings support supervised use of LLMs and evaluation approaches that assess reasoning and prioritisation as well as target detection. They should not be interpreted as evidence that one model is clinically superior in real-world practice. Clinical trial not applicable.

Read PDF

Similar papers

Review Open access Sep 2026

Evaluating large language models as clinical decision support tools in primary healthcare settings: Protocol for a multi-country comparative validation study on expert-adjudicated hypothetical vignettes (hypMOOVE-PHC)

The hypMOOVE-PHC study is the hypothetical vignette phase of the Massive Open Online Validation and Evaluation (MOOVE) initiative, implemented in Kenya, Malawi, and Tanzania, and aims to validate a pool of LLMs through clinical review of expert-generated vignettes through fully crossed repeated-measures comparative eva...

P. Macharia, C. Kachimanga, M. Mahende et al. · 0 citations
Open access Sep 2026

Large language models for late-life depression: a blinded benchmark of clinical safety, geriatric appropriateness, and triage

Background General-purpose large language models are increasingly used by patients and caregivers to obtain mental health information and guidance about when professional care is required. In late-life depression, broadly accurate information may nevertheless be unsafe when cognitive change, multimorbidity, frailty, po...

Wei Xiao, Huan Zhang, Xiao-Yi Chen et al. · 0 citations
#artificial intelligence Review Oct 2026

A rubric landscape for evaluating clinical reasoning in large language models: what exists, what is missing, and what needs to be combined

Exam-style accuracy does not establish whether large language models (LLMs) reason well over clinical records. We define clinical reasoning as integrating and updating evidence across time and sources to form, revise and justify a patient's problem representation and a defensible plan. This structured narrative review...

Zhang Jiang, Zina Ibrahim, J. Teo · 0 citations
Review Open access Sep 2026

LARGE LANGUAGE MODELS IN MEDICATION MANAGEMENT: CLINICAL UTILITY, SAFETY ASSURANCE, AND PHARMACIST-LED GOVERNANCE

Large language models (LLMs) can produce fluent medication explanations, draft patient-facing materials, summarize records, and assist with information retrieval. Their linguistic competence has encouraged proposals for pharmaceutical-care use, yet fluency is not equivalent to clinical reliability. This integrative nar...

Júlia Costa Oliveira Ornelas · 0 citations
Open access Sep 2026

A structured exploratory multidisciplinary evaluation of four large language models responding to frequently asked patient questions about pregabalin

Large Language Models (LLMs) are increasingly consulted by millions of patients seeking pharmaceutical information, yet their reliability and safety in providing medication-related advice remain inadequately evaluated. This study assessed the performance of four leading LLMs in responding to frequently asked questions...

Y. L. T. Bayala, I. A. Tinni, Thierry Boris Wend-Yam Yaméogo et al. · 0 citations
#artificial intelligence Preprint Sep 2026

A Proposed Rubric for Evaluating Expressed Clinical Reasoning in Large Language Model Responses

We propose a rubric for assessing expressed clinical reasoning in model responses, drawing on three bodies of work: medical education assessment frameworks (ART, SCT, Key Feature Problems and OSCE); clinical LLM benchmarks (MedR-Bench, HealthBench, TIMER-Bench, DR. BENCH, PrIME-LLM and PatientSafeBench); and general LL...

Zhang Jiang, Zina Ibrahim, J. Teo · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.