Skip to content
Review

AI-Powered Prescription Error Detection Using Large Language Models (LLMs): A Systematic Review and Future Perspectives

Aug 2026 · International Scientific Journal of Engineering and Management · Vol 05, pp. 1-9 · 0 citations

TL;DR

It is concluded that LLM-based decision-support tools hold substantial promise as complementary — rather than autonomous — decision-support systems capable of transforming medication safety and pharmacy practice.

Abstract

Abstract Medication errors remain among the most significant preventable causes of patient harm worldwide, contributing to increased morbidity, mortality, prolonged hospitalization, and escalating healthcare expenditure. Large Language Models (LLMs) — including GPT-4, Gemini, Claude, and Llama — have emerged as promising clinical decision-support tools capable of interpreting complex medical terminology, analyzing prescriptions in real time, and flagging potential errors before medications reach the patient. This review synthesizes current evidence on the application of LLMs in prescription error detection, presents a consolidated system architecture and operational workflow for LLM-enabled medication safety pipelines, and critically examines their benefits, limitations, and future trajectory. Evidence from recent clinical evaluations indicates that LLM-based decision-support tools can achieve high concordance with expert pharmacist judgment and measurably reduce near-miss medication events when deployed with appropriate safeguards. However, challenges including AI hallucination, data privacy, algorithmic bias, regulatory ambiguity, and the continued necessity of human oversight must be addressed before widespread clinical adoption. The review concludes that LLMs hold substantial promise as complementary — rather than autonomous — decision-support systems capable of transforming medication safety and pharmacy practice. Keywords: Large Language Models; Prescription Error Detection; Medication Safety; Clinical Decision Support; Artificial Intelligence in Healthcare; Electronic Health Records; Pharmacovigilance

View source

Similar papers

Review Open access Aug 2026

Evaluating large language model performance in US FDA regulatory science

This study compares the performance of three LLMs in extracting, analyzing, and synthesizing regulatory and clinical information from FDA drug reviews, guidance for the industry, and drug labels as accessed through their standard user interfaces, using antibiotics approved for complicated urinary tract infections between 2010 and 2025.

Khulud Bukhari, Rosa Rodriguez-Monguio, B. Lopez-Bermudez et al. · 0 citations
Review Open access 2025

Large Language Models in Healthcare: Opportunities and Ethical Challenges

A thorough review of the developments in LLM technologies, their uses in clinical and administrative settings, as well as their ethical considerations are reviewed to suggest a conceptual structure for responsible implementation that will ensure both technological innovation and patient safety, as well as regulatory compliance and ethical health care practices.

Noah Wright · 0 citations
Review Jul 2026

Can AI assist in reducing diagnostic error? A narrative review

It is concluded that AI tools have matured to the extent that they can improve diagnostic decision-making of clinicians and can assist institutions in increasing diagnostic safety.

Ian A. Scott · 0 citations
Review Open access Aug 2026

Large Language Models in Adverse Drug Reaction Detection and Pharmacovigilance: A Systematic Review of Current Applications, Challenges, and Future Directions

Background/Objectives: Pharmacovigilance workflows rely heavily on unstructured text across diverse sources. Here, we systematically reviewed how large language models (LLMs) are being explored as support tools for adverse drug reaction (ADR) detection, extraction, triage, and documentation, highlighting their potential for precision medicine and big data-enabled safety monitoring. Methods: Following the Preferred Reporting Items for Systematic Reviews and Meta-Analyses 2020 guidelines, we systematically searched PubMed, Scopus, and Web of Science for studies published between January 2022 and March 2026. Ultimately, 83 empirical studies satisfied the inclusion criteria. A narrative synthesis was conducted to address methodological heterogeneity across these studies. Results: LLM applications were concentrated in constrained information-extraction and classification tasks, including signal evaluation, clinical-note extraction, social media surveillance, and literature screening. Quantitative performance varied substantially by system design: error-correction prompting yielded an F1-score of 0.921 for ADR named entity recognition, whereas retrieval-augmented generation improved data-retrieval accuracy from 8.3% to 78.3%. Most studies were retrospective, benchmark-based, or proof-of-concept evaluations. Across 581 paired pre-consensus domain judgements, observed inter-rater agreement was 90.4% and Cohen’s κ was 0.837 (95% CI 0.772–0.895). Hallucination, low specificity, prompt sensitivity, narrow datasets, and weak external validation remained common limitations. Conclusions: Current evidence supports supervised, task-specific applications of LLMs for extraction, triage, retrieval, and documentation rather than autonomous pharmacovigilance decision-making. Prospective evaluation, external validation, transparent reporting, and accountable human oversight are required before high-stakes clinical or regulatory deployment.

Tae You Kim, Won-Sik Oh, Dong-Hwa Jeong · 0 citations