Aug 2026· Journal of Studies on Alcohol and Drugs· 0 citations
Medicine
TL;DR
LLMs could serve as a quality control check during opioid policy surveillance research, supplementing human review, and benefit from best practices and technical guidelines for LLM utilization.
Abstract
Objective
Policy surveillance typically involves detailed, time-consuming manual screening of policies for inclusion in a final dataset. This screening process involves risks of human error and inconsistent application of inclusion/exclusion criteria, especially in complicated legal landscapes like the US opioid treatment landscape. Large language models (LLMs) could assist human subject matter experts (SMEs) during screening, but LLMs have been understudied for policy surveillance. Therefore, we conducted a test comparing opioid treatment policy screening decisions between SMEs and an LLM.
Methods
Using a Boolean search string in legal software, we identified 99 potentially relevant Massachusetts policies for emergency department opioid addiction treatment, and we downloaded text from government websites. Next, we compared two approaches to screening those policies using pre-defined inclusion and exclusion criteria: (a) manual screening by three SMEs, and (b) an LLM approach. We assessed the overall percentage of inclusion/exclusion decisions where the LLM made the same decision as the SMEs. We also identified the percentage of policies selected for inclusion by the SMEs with which the LLM agreed and potential reasons for discrepancies.
Results
The LLM made the same decision for 96 of 99 policies (97% of the time). All policies that SMEs chose to include (n=2) were also included by the LLM. Discrepancies reflected implicit inclusion and exclusion criteria used by SMEs but not provided to LLMs.
Conclusion
LLMs could serve as a quality control check during opioid policy surveillance research, supplementing human review. The policy surveillance field would benefit from best practices and technical guidelines for LLM utilization.
To develop and validate ScreenAgent, a large language model (LLM) agent for title and abstract screening, and a review-specific method for prospectively estimating screening performance, which identified nearly all eligible studies with human-level reliability for a fraction of a US cent per record while keeping human reviewers as the final arbiters.
D. Dobin, A. Witmer, F. Sweeney et al.· medRxiv· 0 citations
It is concluded that LLM-based decision-support tools hold substantial promise as complementary — rather than autonomous — decision-support systems capable of transforming medication safety and pharmacy practice.
K. K. Kumar, Koyya Gowtham Reddy, K. Reddy· International Scientific Jou...· 0 citations
Background/Objectives: Pharmacovigilance workflows rely heavily on unstructured text across diverse sources. Here, we systematically reviewed how large language models (LLMs) are being explored as support tools for adverse drug reaction (ADR) detection, extraction, triage, and documentation, highlighting their potential for precision medicine and big data-enabled safety monitoring. Methods: Following the Preferred Reporting Items for Systematic Reviews and Meta-Analyses 2020 guidelines, we systematically searched PubMed, Scopus, and Web of Science for studies published between January 2022 and March 2026. Ultimately, 83 empirical studies satisfied the inclusion criteria. A narrative synthesis was conducted to address methodological heterogeneity across these studies. Results: LLM applications were concentrated in constrained information-extraction and classification tasks, including signal evaluation, clinical-note extraction, social media surveillance, and literature screening. Quantitative performance varied substantially by system design: error-correction prompting yielded an F1-score of 0.921 for ADR named entity recognition, whereas retrieval-augmented generation improved data-retrieval accuracy from 8.3% to 78.3%. Most studies were retrospective, benchmark-based, or proof-of-concept evaluations. Across 581 paired pre-consensus domain judgements, observed inter-rater agreement was 90.4% and Cohen’s κ was 0.837 (95% CI 0.772–0.895). Hallucination, low specificity, prompt sensitivity, narrow datasets, and weak external validation remained common limitations. Conclusions: Current evidence supports supervised, task-specific applications of LLMs for extraction, triage, retrieval, and documentation rather than autonomous pharmacovigilance decision-making. Prospective evaluation, external validation, transparent reporting, and accountable human oversight are required before high-stakes clinical or regulatory deployment.
Tae You Kim, Won-Sik Oh, Dong-Hwa Jeong· Diagnostics· 0 citations
BACKGROUND
We evaluated the reliability of large language models (LLMs) for abstract screening under real-world review practices and qualitatively characterized model-human discordance to inform safe workflow integration.
METHODS
We evaluated GPT-4.0, GPT-5.0, and GPT-5.0-mini on two curated systematic review datasets representing contrasting topic densities, defined by the target-to-background ratio (TBR): a core-subject dataset (TBR 69%) in which the target intervention was central, and a peripheral-subject dataset (TBR 4.4%) in which the target intervention was incidental. Final inclusion after full-text review served as the reference standard. We developed a qualitative taxonomy of disagreements, classifying false negatives as intended human leniency, gray-zone ambiguity, or true LLM misses, and false positives as implicit or additional human exclusion rules, gray-zone ambiguity, or nominal inclusions that increase workload only.
RESULTS
GPT-5.0-mini achieved the best sensitivity-efficiency trade-off (core-subject: 91% sensitivity with 96.7% workload reduction; peripheral-subject: 83% sensitivity with 92.7% workload reduction) and negative predictive value >99% in both datasets. Disagreement was lower when relevance was central (core-subject: 1.6%, 7/430) with no true LLM misses (0/430). In the peripheral-subject dataset, disagreement was higher (10.6%, 74/696), driven mainly by intended human leniency among false negatives (52/56) and gray-zone ambiguity among false positives (12/18), while true LLM misses remained rare (0.4%, 3/696).
CONCLUSION
Many model-human disagreements reflect topic- and workflow-dependent screening conventions rather than intrinsic model failure. LLM-assisted screening may improve efficiency without compromising reliability when accompanied by appropriate safeguards for ambiguous records.
K. Lee, Hakyoung Kim, Dae Sik Yang et al.· Journal of Epidemiology· 0 citations
These findings identify unsupported reassurance as measurable evidence-boundary behavior in LLM drug-risk assessment and establish a reproducible framework for auditing adherence to an externally supplied evidence boundary, defined by PubMed evidence classification and enforced by prompt policy rather than inferred independently by the model.
Jinwan Shi, Yong Ma, Yinhui Liu et al.· Frontiers in Public Health· 0 citations
: Large language models can generate fluent and often convincing answers in medical contexts, but in high-risk settings, fluency alone is not enough. This paper presents a policy-guided LLM pipeline for medication-related clinical decision support, designed to classify prompts into safety-oriented decision categories: ACCEPT, WARN, DEFER, ESCALATE, or REFUSE. The system combines rule-based risk detection, guardrail routing, and a decision policy layer to handle prompts involving dosage, pregnancy, drug interactions, self-adjustment, and other safety-critical contexts. The pipeline was evaluated on a benchmark of 55 medication-related questions using two baseline models. Both models produced the same overall system accuracy of 60.0%, with a false accept rate of 16.4%, suggesting that the main limitations are not model-specific but structural. Three recurring failure modes emerged: false acceptance of implicitly risky prompts, over-escalation of educational or professional-context questions, and under-escalation of self-adjustment or dangerous-intent cases. These findings make the system useful not only as a prototype, but also as a transparent framework for studying where safety-oriented LLM pipelines succeed and where they still fail.
Ana Stevanović, Mlađan Jovanović· SINTEZA· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduAug 17, 2026
A USAF cadet and a Lincoln Laboratory researcher found AI chatbots can help nontechnical service members produce viable software applications for their unique problems.