Aug 2026· Discover Artificial Intelligence· 0 citations
TL;DR
This study compares the performance of three LLMs in extracting, analyzing, and synthesizing regulatory and clinical information from FDA drug reviews, guidance for the industry, and drug labels as accessed through their standard user interfaces, using antibiotics approved for complicated urinary tract infections between 2010 and 2025.
Abstract
Clinical and population decision-making relies on the systematic evaluation of extensive regulatory evidence. The FDA drug reviews provide detailed information on clinical trial design, enrollment criteria, sample size, randomization, comparators, endpoints, and indications. However, extracting these data is resource-intensive and time-consuming. Generative Artificial Intelligence large language models (LLMs) may accelerate the extraction and synthesis of such information. This study compares the performance of three LLMs, ChatGPT-4o, Gemini 2.5 Pro, and DeepSeek R1, in extracting, analyzing, and synthesizing regulatory and clinical information from FDA drug reviews, guidance for the industry, and drug labels as accessed through their standard user interfaces, using antibiotics approved for complicated urinary tract infections (cUTIs) between 2010 and 2025.
LLMs were evaluated using general (short, direct) and detailed (structured, guidance-referencing) prompts across five domains including accuracy (precision and recall), explanation quality, error type (hallucination rate, misclassification, and omission), operational efficiency (response time, correct answers per second, and seconds per correct answer), and consistency with responses generated in duplicate runs. Two investigators independently reviewed outputs against FDA guidance, resolving discrepancies by consensus. Statistical analyses included χ
2
, Wilcoxon, and Kruskal–Wallis tests with false discovery rate correction. Mixed-effects logistic regression was conducted to account for clustering by drug and question.
Among 324 responses, accuracy differed significantly across models (χ
2
,
p
< 0.001) with Gemini 2.5 Pro achieving the highest accuracy (66.7%), followed by ChatGPT-4o (51.9%) and DeepSeek R1 (37.0%). General prompts were associated with higher accuracy than detailed prompts (59.3% vs 44.4%;
p
= 0.011). Gemini 2.5 Pro showed the highest explanation quality, while Gemini 2.5 Pro and ChatGPT-4o showed comparable consistency and DeepSeek R1 was less consistent. Hallucination was the most frequent error type across models.
LLMs showed variable capability in extracting regulatory and clinical information. Gemini 2.5 Pro showed the strongest overall performance, while ChatGPT-4o was faster but less accurate, and DeepSeek R1 underperformed across most domains. These findings highlight both the promise and limitations of LLMs in regulatory science and support their complementary use with human review to support evidence extraction and synthesis for regulatory review.
INTRODUCTION
Traditional pharmacovigilance relies on slow clinical trials and post-marketing studies with limited coverage. This review synthesizes evidence on Real-World Data (RWD) integration with Artificial Intelligence (AI) for enhanced Adverse Drug Reaction (ADR) detection, evaluates generative AI like ChatGPT-4 and LLaMA-2 in Substance Use Disorder (SUD) scenarios, discusses current applications, and outlines future directions. The objective is to guide researchers, clinicians, and regulators in this evolving field.
METHODS
Literature was reviewed on RWD sources (EHRs, claims, registries, wearables), AI algorithms (supervised/ unsupervised learning, NLP, deep learning), and regulatory frameworks. Generative AI performance was assessed via clinician-blind evaluation of responses to Reddit-sourced SUD queries from r/stopdrinking, r/leaves, and r/OpiatesRecovery, with fact-checking against SAMHSA/FDA guidelines and consistency testing. Data included tables comparing RWD, algorithms, and AI models.
RESULTS
AI enables real-time ADR signals via RWD-AI in CCM, improving diagnostics, personalization, and drug discovery. ChatGPT-4 suggested unsafe opioid microdosing; LLaMA-2 referenced nonexistent resources and improper Xanax sharing, both showing severe inaccuracies in SUD contexts. Tables highlight RWD applications, algorithm uses, and AI limitations like bias and inconsistency.
DISCUSSION
RWD-AI transforms pharmacovigilance but faces bias, transparency, and validation challenges. FHIR/DLT enhance secure exchange; generative AIs require oversight. Implications include equitable safety monitoring via bias mitigation and regulatory compliance.
CONCLUSION
AI-RWD integration advances ADR detection and personalized safety, despite generative AI risks in SUD management. Future success demands validated LLMs, FHIR/blockchain infrastructure, and clinician collaboration for comprehensive, equitable pharmacovigilance.
Pritam Kayal, Priya Manna, Ramit Rahaman et al.· Current pharmaceutical desig...· 0 citations
It is concluded that LLM-based decision-support tools hold substantial promise as complementary — rather than autonomous — decision-support systems capable of transforming medication safety and pharmacy practice.
K. K. Kumar, Koyya Gowtham Reddy, K. Reddy· International Scientific Jou...· 0 citations
Findings show that clinical LLM explainability has shifted toward fluent generative rationales, but evidence that such explanations reflect model reasoning remains limited, and three regulatory priorities are highlighted: prioritizing explanations that enable independent verification or logic auditing over plausibility-only rationales; preferring inspectable models where regulatory documentation is required; and prospectively validating explanations in clinical workflows before scaling.
Marzouk et al. reviewed 147 studies on artificial intelligence (AI) applications for predicting drug–drug, drug–disease, and drug–nutrient interactions, providing a broad overview of current machine learning and deep-learning approaches. However, several methodological and conceptual limitations reduce the reproducibility and interpretability of the review. The search strategy appears largely restricted to PubMed with title- and abstract-level filtering, while manual record removal is reported without explicit criteria defining “irrelevant” studies, limiting transparency and reproducibility. Protocol registration, duplicate independent screening, standardized extraction procedures, and formal bias assessment using established frameworks such as ROBIS, PROBAST+AI, and TRIPOD+AI were not clearly reported. The review reports performance metrics such as area under the receiver operating characteristic curve (AUROC), but does not provide a structured framework for interpreting or comparing metrics across heterogeneous datasets, prediction tasks, and evaluation protocols. Because the interpretation of AUROC and precision–recall metrics depends on class prevalence, outcome definition, and the intended prediction task, future reviews should report complementary discrimination metrics, calibration, uncertainty estimates, and external validation rather than assuming that any single metric is universally preferable. Claims of superior model performance should be supported by confidence intervals and statistical comparisons appropriate to the evaluation design, such as paired DeLong testing when applicable. Claims of superior model performance should also be supported by appropriate statistical testing, including methods such as the nonparametric DeLong test. Several conceptual clarifications are also warranted. AI models may prioritize hypotheses but do not replace experimental or clinical validation under current regulatory standards. Furthermore, AUROC should not be conflated with pharmacokinetic area under the curve, and SciBERT should not be characterized as a three-dimensional molecular graph framework. Future reviews should adopt transparent multi-database searches, structured bias assessment, and reproducible reporting practices.
Alireza Kargar, Mohammad Ali Zamani, Ghader Mohammadnezhad· Journal of Cheminformatics· 0 citations
Background/Objectives: Pharmacovigilance workflows rely heavily on unstructured text across diverse sources. Here, we systematically reviewed how large language models (LLMs) are being explored as support tools for adverse drug reaction (ADR) detection, extraction, triage, and documentation, highlighting their potential for precision medicine and big data-enabled safety monitoring. Methods: Following the Preferred Reporting Items for Systematic Reviews and Meta-Analyses 2020 guidelines, we systematically searched PubMed, Scopus, and Web of Science for studies published between January 2022 and March 2026. Ultimately, 83 empirical studies satisfied the inclusion criteria. A narrative synthesis was conducted to address methodological heterogeneity across these studies. Results: LLM applications were concentrated in constrained information-extraction and classification tasks, including signal evaluation, clinical-note extraction, social media surveillance, and literature screening. Quantitative performance varied substantially by system design: error-correction prompting yielded an F1-score of 0.921 for ADR named entity recognition, whereas retrieval-augmented generation improved data-retrieval accuracy from 8.3% to 78.3%. Most studies were retrospective, benchmark-based, or proof-of-concept evaluations. Across 581 paired pre-consensus domain judgements, observed inter-rater agreement was 90.4% and Cohen’s κ was 0.837 (95% CI 0.772–0.895). Hallucination, low specificity, prompt sensitivity, narrow datasets, and weak external validation remained common limitations. Conclusions: Current evidence supports supervised, task-specific applications of LLMs for extraction, triage, retrieval, and documentation rather than autonomous pharmacovigilance decision-making. Prospective evaluation, external validation, transparent reporting, and accountable human oversight are required before high-stakes clinical or regulatory deployment.
Tae You Kim, Won-Sik Oh, Dong-Hwa Jeong· Diagnostics· 0 citations
The authors' analysis reveals that LLMs demonstrate promising capabilities in processing textual and visual data related to various liver diseases, including hepatocellular carcinoma, cirrhosis, and non-alcoholic fatty liver disease, but study heterogeneity and significant challenges remain regarding accuracy, reliability, and safety.
T. Suenghataiphorn, Narisara Tribuddharat, Pojsakorn Danpanichkul et al.· Hepatology Forum· 0 citations