Skip to content
Review Open access

Making Broad Evidence Synthesis Feasible: An LLM Screening Agent for Meta-Analyses Applied To Suicide Prevention

Aug 2026 · medRxiv · 0 citations
Medicine

TL;DR

To develop and validate ScreenAgent, a large language model (LLM) agent for title and abstract screening, and a review-specific method for prospectively estimating screening performance, which identified nearly all eligible studies with human-level reliability for a fraction of a US cent per record while keeping human reviewers as the final arbiters.

Abstract

Importance. Systematic reviews and meta-analyses inform suicide-prevention policy and practice, but broad database searches are difficult to screen manually. This limits capture of upstream interventions, such as economic policies, with indirect effects on suicide. Reliable automated screening could make broader and more comprehensive evidence syntheses feasible. Objective. To develop and validate ScreenAgent, a large language model (LLM) agent for title and abstract screening, and a review-specific method for prospectively estimating screening performance. Design, Setting, and Participants. ScreenAgent was validated internally on a prospective meta-analysis, and externally on two published systematic reviews. The correct include and exclude decisions followed standard systematic-review screening methodology. Exposures. ScreenAgent, an LLM agent returning structured include-or-exclude decisions. Records it marked for inclusion were re-checked by a second, cascade pass using a higher-effort LLM. For the external reviews, the agent's prompt was tuned automatically on a small set of labeled examples. Main Outcomes and Measures. We calculated sensitivity, specificity, workload reduction (the percentage of records removed from human review), and agent-versus-human reliability via Cohen kappa. Sensitivity was estimated by direct comparison (internal) and 5-fold cross-validation (external). Results. In the internal validation, ScreenAgent identified 43 of 44 eligible studies (sensitivity 97.7%; 95% CI, 88.2%-99.6%) with a generic prompt applied without any review-specific optimization, specificity 98.0%, and a measured full-corpus workload reduction of 99.4%. The cost was $855.91 for the full 201,064-record corpus (0.43 US cents per record). Agent-versus-human-consensus agreement exceeded human-versus-human agreement (Cohen kappa 0.75 vs 0.64; percent agreement 97.3% vs 95.4%). For two external validation studies, automatic tuning resulted in a cross-validated sensitivity of 95.9% (95% CI, 90.0%-98.4%) and 97.4% (90.9%-99.3%), with workload reductions of 97.4% and 98.4%. Conclusions and Relevance. Suicide prevention efforts often require rapid consolidation of evidence because of the inherent challenges of single studies trying to prevent rare outcomes. On both internal and external validation sets, ScreenAgent identified nearly all eligible studies with human-level reliability for a fraction of a US cent per record while keeping human reviewers as the final arbiters. By making broad searches feasible and screening performance measurable beforehand, this approach can serve as a transparent methodology to strengthen the speed at which we can inform and advance suicide prevention efforts.

Read PDF

Similar papers

Review Open access Jul 2026

Scoping evidence review on the prevention and treatment of opioid overdose.

BACKGROUND AND AIM Opioid overdose is a public health crisis. Globally, ~ 600,000 deaths each year are attributable to drug use, and roughly three-quarters are linked to opioids. The rising prevalence of high-potency synthetic opioids like fentanyl necessitates comprehensive prevention and treatment strategies. This scoping review aims to map the published literature on the early detection and community management of opioid overdose and post-overdose care and identify research gaps. The aim was also to inform the potential research questions for updating the WHO and other international guidelines. METHODS We systematically searched PubMed, EMBASE, PROSPERO, and the Cochrane Database for systematic reviews for studies published between January 2014 and March 2024. Inclusion criteria encompassed systematic reviews and review protocols, randomized clinical trials, and their secondary analysis, and guidelines published in those databases, addressing opioid overdose early detection and treatment strategies. RESULTS A total of 3,057 unique records were screened, resulting in 105 studies meeting the inclusion criteria. Thirty studies examined the effects of the acute opioid overdose treatment. Among the preventive strategies, 46 studies focused on "Take Home Naloxone" strategies, reflecting a significant emphasis on naloxone distribution as a primary intervention. However, 11 studies examined post-overdose care, highlighting a notable gap in the literature. Additionally, research predominantly concentrated in the Americas (70 studies) and Europe (27 studies), with limited representation from Africa and Southeast Asia. There was a marked absence of studies on polysubstance use and ultrapotent opioids, with one systematic review addressing the treatment of fentanyl overdose. CONCLUSION Findings support future international guidelines, such as those of the WHO, to examine the effectiveness of take-home naloxone, peer training, and community-based distribution models, higher-dose naloxone protocols for ultrapotent opioids, and digital delivery models. This review identifies gaps in post-overdose care, fentanyl and polysubstance overdose management, and research from low-resource settings.

Abhishek Ghosh, Shinjini Choudhury, Tathagata Mahintamani et al. · 0 citations
Review Open access Aug 2026

Uncovering Means Restriction Activities for Suicide Prevention That Have Been Undertaken in Australia: A Scoping Review.

OBJECTIVES To map and summarize evidence on means restriction activities that have been implemented for the prevention of suicide across Australia. Further, to provide an overview of the impact of the activities on suicide rates and ascertain barriers/challenges, costs, and any other relevant associated information. METHODS We conducted a scoping review of the literature following JBI-formerly known as Joanna Briggs Institute-methodology for scoping reviews and recommendations by Cochrane Rapid Reviews. Four databases were searched for identifying all published material in English. Gray literature was searched through Google Scholar and state-based and national health commission websites. RESULTS We identified five publications addressing means restriction activities implemented for suicide prevention. All studies describe restriction of access to places where there had been suicides by jumping from a height. The studies concluded that mean restriction activities were valuable in the prevention of suicide by jumping. Further, we reviewed four additional publications describing activities that restricted access to means (e.g., firearms legislation) but where the intervention was not aimed at suicide prevention per se, albeit suicide rates were reported. CONCLUSION Despite there being few studies available, our review identifies evidence on means restriction activities for the prevention of suicide implemented within Australia and provides a summary of implemented interventions that have been reported to be associated with a decrease in suicide rates. In particular, the apparent benefit of structural interventions at jumping sites supports their broader adoption: any such future activities require careful evaluation.

Julia Teng Jia Ru, David Castle, Myles N Moore et al. · 0 citations
Review Open access Aug 2026

Cancer prevention and screening for people with intellectual disabilities in Europe: a systematic review and recommendations for policy and practice

People with intellectual disability experience inequalities across the cancer care continuum, including lower screening participation, later-stage diagnosis, and higher cancer related mortality. Whether European cancer prevention and screening policies explicitly address this population has not been systematically mapped. This review aimed to identify and synthesise European national and international policy frameworks, guidelines, and recommendations for cancer prevention and screening that address this population. A systematic search of policy, recommendation and guideline documents across academic databases, grey literature portals, and organisational sources, to August 2025 was under taken. The Population, Concept, Context (PCC) framework guided eligibility. Two reviewers per country screened, extracted, and appraised documents indepen dently using the Authority, Accuracy, Coverage, Objectivity, Date, Significance (AACODS) checklist. Synthesis was narrative, stratified by cancer site and cancer control focus. Sixteen documents met the inclusion criteria, originating from five jurisdictions: the United Kingdom, the Netherlands, Belgium, Denmark, and the European Union more broadly. Fifteen of the sixteen addressed cancer screen ing, one also addressed prevention, and two addressed clinical monitoring. Adjustments for people with intellectual disability were concentrated on breast, cervical, and bowel screening and grouped into accessible information, appoint ment adaptations, carer involvement, consent and capacity support, and work force training. Most documents originated from higher expenditure healthcare systems, and no eligible documents were retrieved from the majority of European Union member states. Policy guidance for cancer prevention and screening for people with intellectual disability is unevenly distributed across Europe and con centrated in a small group of countries. Existing guidance should be extended beyond screening to encompass prevention, treatment, and survivorship, and EU coordination may help address this gap.

V. Vuković, Peter Knapp, Kate Sykes et al. · 0 citations
#software testing Review Open access Aug 2026

Proof of Concept of Large Language Models for Opioid Treatment Policy Surveillance: 97% Agreement With Subject Matter Experts.

LLMs could serve as a quality control check during opioid policy surveillance research, supplementing human review, and benefit from best practices and technical guidelines for LLM utilization.

B. Andraka-Christou, Jae Park, F. Ahmed et al. · 0 citations
Review Open access Jul 2026

Comparative effects of drugs for adults with overweight or obesity: systematic review and network meta-analysis

Abstract Objective To provide an up-to-date evidence summary about the comparative benefits and harms of drugs for adults with overweight or obesity to inform decision making for policymakers, payers, clinicians, and patients. Design Systematic review and network meta-analysis of 24 outcomes using frequentist random effects models and bayesian dose-response models, the GRADE (Grading of Recommendations Assessment, Development, and Evaluation) approach, and the Cochrane Risk of Bias 2 tool. Data sources Medline, Embase, and Cochrane Library, searched up to 12 November 2025. Study selection Randomised controlled trials of 12 weeks’ duration or longer comparing one or more drugs with lifestyle modification, placebo, or another drug. Results This network meta-analysis comprised 262 trials (99 791 participants) evaluating 19 drugs with follow-up from 12 to 172 weeks. Compared with lifestyle modification alone, at one year, moderate to high certainty evidence shows substantial weight loss with tirzepatide (mean difference -14.9%, 95% confidence interval -16.0% to -13.9%), cagrilintide-semaglutide (CagriSema, -14.8%, -16.9% to -12.7%), oral semaglutide (-10.9%, -12.7% to -9.1%), orforglipron (-9.9%, -12.4% to -7.5%), subcutaneous semaglutide (-9.8%, -10.6% to -9.1%), and phentermine-topiramate (-8.1%, -9.7% to -6.5%). Emerging agents (ecnoglutide, mazdutide, retatrutide) may produce similar or greater reductions (13.1-14.6%; very low to low certainty). Moderate to high certainty evidence supports discontinuation because of adverse events to be highest with orforglipron, naltrexone-bupropion, liraglutide, phentermine-topiramate, CagriSema, and oral semaglutide (risk ratios from 1.9 to 4.2); gastrointestinal events were most increased with naltrexone-bupropion, oral semaglutide, orforglipron, and tirzepatide (risk ratios from 3.1 to 4.2). Fatigue risk increased, particularly with naltrexone-bupropion (risk ratio 8.9; absolute increase 331 per 1000 people over one year), orforglipron (3.4; 100 more per 1000), and CagriSema (3.2; 92 more per 1000). Tirzepatide reduced fat mass the most (by 25.7%) but also lean mass the most (by 8.3%). Subcutaneous semaglutide was the only drug associated with reduced all cause mortality (risk ratio 0.81, 95% confidence interval 0.72 to 0.93) and myocardial infarction (0.72, 0.61 to 0.85), with these estimates largely informed by cardiovascular outcome trials in high risk populations. Subcutaneous semaglutide (0.43, 0.21 to 0.84) and tirzepatide (0.49, 0.27 to 0.88) reduced heart failure risk. No drugs convincingly reduced kidney failure or improved quality of life (43 trials with 45 663 participants) beyond established minimally important differences (all mean differences <5 points; minimally important difference 10). Except for larger weight reductions in trials with longer duration (shown for subcutaneous semaglutide), subgroup analyses for drug dosages and key patient characteristics did not identify credible differences in relative effects of treatment. Conclusions Obesity drugs produce variable weight loss at one year, with larger benefits generally accompanied by greater harms and discontinuation. Most agents do not improve quality of life meaningfully and few show cardiovascular benefits. Decisions in clinical practice should consider trade-offs between benefits and harms within the context of shared decision making. Systematic review registration PROSPERO CRD42024507993.

K. Nong, Qingyang Shi, Xinran Xie et al. · 3 citations
Review Open access 2026

The use of AI scribes across medical fields and why psychiatry is lagging behind: A scoping review.

Objective This scoping review aimed to examine the current available literature on the use of artificial intelligence (AI) scribes across all medical fields compared to their use in psychiatry. Methods A scoping review was conducted following the Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews (PRISMA-ScR) guidelines. Five electronic databases (PubMed, MEDLINE, EMBASE, PsycINFO, and CINAHL) were systematically searched for studies published up to June 2025. Two reviewers independently screened titles and abstracts according to pre-established eligibility criteria using the PICOS framework. Full-text articles were then assessed for inclusion. Results A total of 1896 records were identified, of which 15 met the inclusion criteria. Most studies were conducted in outpatient, primary care, and procedural settings. Across studies, clinicians generally reported favorable acceptance and perceptions of AI scribes. Findings also suggested improved clinician well-being, reduced documentation time, better workflow efficiency, and positive patient experience. Results on documentation accuracy and quality were mixed. Critically, no eligible studies evaluated the use of AI scribes in psychiatry. Limitations included study heterogeneity, with wide variations in study design, sample size, and evaluation metrics. Conclusion The absence of empirical studies in psychiatry highlights a significant gap in the literature, particularly in contrast to the growing body of research across other medical specialties. Because of its unique clinical, ethical, and relational dimensions, dedicated research is needed to evaluate the feasibility, safety, and ethical implications of implementing AI scribes in psychiatric care.

Chloé Daignault, Justine Audelin-Rinfret, Marie Désilets et al. · 0 citations