Quality and safety of large language model–generated medication review outputs in geriatric pharmacotherapy: a two-stage comparative vignette-based benchmark evaluation
These findings support supervised use of large language models and evaluation approaches that assess reasoning and prioritisation as well as target detection and should not be interpreted as evidence that one model is clinically superior in real-world practice.
Kubra Cingar Alpay, D. Ozata, T. Gedik et al.
· BMC Geriatrics · 0 citations