Skip to content
#software testing Review Open access

Proof of Concept of Large Language Models for Opioid Treatment Policy Surveillance: 97% Agreement With Subject Matter Experts.

Aug 2026 · Journal of Studies on Alcohol and Drugs · 0 citations
Medicine

TL;DR

LLMs could serve as a quality control check during opioid policy surveillance research, supplementing human review, and benefit from best practices and technical guidelines for LLM utilization.

Abstract

Objective

Policy surveillance typically involves detailed, time-consuming manual screening of policies for inclusion in a final dataset. This screening process involves risks of human error and inconsistent application of inclusion/exclusion criteria, especially in complicated legal landscapes like the US opioid treatment landscape. Large language models (LLMs) could assist human subject matter experts (SMEs) during screening, but LLMs have been understudied for policy surveillance. Therefore, we conducted a test comparing opioid treatment policy screening decisions between SMEs and an LLM.

Methods

Using a Boolean search string in legal software, we identified 99 potentially relevant Massachusetts policies for emergency department opioid addiction treatment, and we downloaded text from government websites. Next, we compared two approaches to screening those policies using pre-defined inclusion and exclusion criteria: (a) manual screening by three SMEs, and (b) an LLM approach. We assessed the overall percentage of inclusion/exclusion decisions where the LLM made the same decision as the SMEs. We also identified the percentage of policies selected for inclusion by the SMEs with which the LLM agreed and potential reasons for discrepancies.

Results

The LLM made the same decision for 96 of 99 policies (97% of the time). All policies that SMEs chose to include (n=2) were also included by the LLM. Discrepancies reflected implicit inclusion and exclusion criteria used by SMEs but not provided to LLMs.

Conclusion

LLMs could serve as a quality control check during opioid policy surveillance research, supplementing human review. The policy surveillance field would benefit from best practices and technical guidelines for LLM utilization.

Read PDF

Similar papers

Review Open access Aug 2026

Making Broad Evidence Synthesis Feasible: An LLM Screening Agent for Meta-Analyses Applied To Suicide Prevention

To develop and validate ScreenAgent, a large language model (LLM) agent for title and abstract screening, and a review-specific method for prospectively estimating screening performance, which identified nearly all eligible studies with human-level reliability for a fraction of a US cent per record while keeping human reviewers as the final arbiters.

D. Dobin, A. Witmer, F. Sweeney et al. · 0 citations
Review Aug 2026

AI-Powered Prescription Error Detection Using Large Language Models (LLMs): A Systematic Review and Future Perspectives

It is concluded that LLM-based decision-support tools hold substantial promise as complementary — rather than autonomous — decision-support systems capable of transforming medication safety and pharmacy practice.

K. K. Kumar, Koyya Gowtham Reddy, K. Reddy · 0 citations
Review Open access Aug 2026

Large Language Models in Adverse Drug Reaction Detection and Pharmacovigilance: A Systematic Review of Current Applications, Challenges, and Future Directions

Background/Objectives: Pharmacovigilance workflows rely heavily on unstructured text across diverse sources. Here, we systematically reviewed how large language models (LLMs) are being explored as support tools for adverse drug reaction (ADR) detection, extraction, triage, and documentation, highlighting their potential for precision medicine and big data-enabled safety monitoring. Methods: Following the Preferred Reporting Items for Systematic Reviews and Meta-Analyses 2020 guidelines, we systematically searched PubMed, Scopus, and Web of Science for studies published between January 2022 and March 2026. Ultimately, 83 empirical studies satisfied the inclusion criteria. A narrative synthesis was conducted to address methodological heterogeneity across these studies. Results: LLM applications were concentrated in constrained information-extraction and classification tasks, including signal evaluation, clinical-note extraction, social media surveillance, and literature screening. Quantitative performance varied substantially by system design: error-correction prompting yielded an F1-score of 0.921 for ADR named entity recognition, whereas retrieval-augmented generation improved data-retrieval accuracy from 8.3% to 78.3%. Most studies were retrospective, benchmark-based, or proof-of-concept evaluations. Across 581 paired pre-consensus domain judgements, observed inter-rater agreement was 90.4% and Cohen’s κ was 0.837 (95% CI 0.772–0.895). Hallucination, low specificity, prompt sensitivity, narrow datasets, and weak external validation remained common limitations. Conclusions: Current evidence supports supervised, task-specific applications of LLMs for extraction, triage, retrieval, and documentation rather than autonomous pharmacovigilance decision-making. Prospective evaluation, external validation, transparent reporting, and accountable human oversight are required before high-stakes clinical or regulatory deployment.

Tae You Kim, Won-Sik Oh, Dong-Hwa Jeong · 0 citations
Review Open access Jul 2026

Qualitative Analysis of Discrepancy Patterns Between Large Language Models and Human Reviewers in Abstract Screening for Systematic Reviews.

BACKGROUND We evaluated the reliability of large language models (LLMs) for abstract screening under real-world review practices and qualitatively characterized model-human discordance to inform safe workflow integration. METHODS We evaluated GPT-4.0, GPT-5.0, and GPT-5.0-mini on two curated systematic review datasets representing contrasting topic densities, defined by the target-to-background ratio (TBR): a core-subject dataset (TBR 69%) in which the target intervention was central, and a peripheral-subject dataset (TBR 4.4%) in which the target intervention was incidental. Final inclusion after full-text review served as the reference standard. We developed a qualitative taxonomy of disagreements, classifying false negatives as intended human leniency, gray-zone ambiguity, or true LLM misses, and false positives as implicit or additional human exclusion rules, gray-zone ambiguity, or nominal inclusions that increase workload only. RESULTS GPT-5.0-mini achieved the best sensitivity-efficiency trade-off (core-subject: 91% sensitivity with 96.7% workload reduction; peripheral-subject: 83% sensitivity with 92.7% workload reduction) and negative predictive value >99% in both datasets. Disagreement was lower when relevance was central (core-subject: 1.6%, 7/430) with no true LLM misses (0/430). In the peripheral-subject dataset, disagreement was higher (10.6%, 74/696), driven mainly by intended human leniency among false negatives (52/56) and gray-zone ambiguity among false positives (12/18), while true LLM misses remained rare (0.4%, 3/696). CONCLUSION Many model-human disagreements reflect topic- and workflow-dependent screening conventions rather than intrinsic model failure. LLM-assisted screening may improve efficiency without compromising reliability when accompanied by appropriate safeguards for ambiguous records.

K. Lee, Hakyoung Kim, Dae Sik Yang et al. · 0 citations
Review Open access Aug 2026

Confident but unsupported: auditing large language models against supplied evidence boundaries in drug-induced liver injury assessment

These findings identify unsupported reassurance as measurable evidence-boundary behavior in LLM drug-risk assessment and establish a reproducible framework for auditing adherence to an externally supplied evidence boundary, defined by PubMed evidence classification and enforced by prompt policy rather than inferred independently by the model.

Jinwan Shi, Yong Ma, Yinhui Liu et al. · 0 citations
Conference Open access 2026

A Policy-guided LLM Pipeline for Safer Clinical Decision Support: Error Analysis on High-risk Medication Queries

: Large language models can generate fluent and often convincing answers in medical contexts, but in high-risk settings, fluency alone is not enough. This paper presents a policy-guided LLM pipeline for medication-related clinical decision support, designed to classify prompts into safety-oriented decision categories: ACCEPT, WARN, DEFER, ESCALATE, or REFUSE. The system combines rule-based risk detection, guardrail routing, and a decision policy layer to handle prompts involving dosage, pregnancy, drug interactions, self-adjustment, and other safety-critical contexts. The pipeline was evaluated on a benchmark of 55 medication-related questions using two baseline models. Both models produced the same overall system accuracy of 60.0%, with a false accept rate of 16.4%, suggesting that the main limitations are not model-specific but structural. Three recurring failure modes emerged: false acceptance of implicitly risky prompts, over-escalation of educational or professional-context questions, and under-escalation of self-adjustment or dangerous-intent cases. These findings make the system useful not only as a prototype, but also as a transparent framework for studying where safety-oriented LLM pipelines succeed and where they still fail.

Ana Stevanović, Mlađan Jovanović · 0 citations

Related blog posts

MIT News · Artificial Intelligence Aug 17, 2026

Q&A: Rethinking how innovation happens

In his latest book, Professor Eugene Fitzgerald examines the forces that turn breakthroughs into value — and why innovation resists simple formulas.