Skip to content
Review

Evaluating human and LLM screening workflows in a conceptually complex scoping review: Recall--workload trade-offs and run-to-run consistency

Aug 2026 · 0 citations · 43 references
Computer Science

TL;DR

Compared human and LLM title-and-abstract screening workflows in a preregistered study embedded in a conceptually complex scoping review, finding that large language models are better suited to validated, auditable, human-supervised workflows than autonomous exclusion.

Abstract

Background. Large language models (LLMs) are increasingly used for screening in evidence synthesis, where false negatives can remove relevant studies before full-text assessment. We compared human and LLM title-and-abstract screening workflows in a preregistered study embedded in a conceptually complex scoping review. Methods. After a conservative title-only screen, 1,131 records were screened by one review lead, four trained assistants screening non-overlapping subsets, and seven complete LLM runs using different models and processing configurations, including a nominally identical repeat run. We compared retained workload, operational recall against 316 verified eligible records, agreement, run-to-run consistency, and procedural burden. Because eligibility was verified only for records advanced and assessed in the parent review, recall estimates were operational. Results. No workflow recovered all verified eligible records. The human workflows and two GPT-5.4 file-batch runs retained 42.2-45.0% of records while achieving 82.3-82.9% recall. Gemini 3.1 file batches achieved the highest recall (83.9%) but retained 56.7% of records. All-at-once configurations recovered fewer eligible records than corresponding file-batch configurations. Two nominally identical GPT-5.4 file-batch runs agreed on 91.7% of records but differed on 94 records, including 29 verified eligible records retained by only one run. Discussion. LLM screening performance depended on the implemented workflow, not model identity alone. Processing configuration, workload, record-level variation, and human-LLM decision integration are therefore substantive properties of deployed systems. For high-recall tasks, LLMs are better suited to validated, auditable, human-supervised workflows than autonomous exclusion.

View source

Similar papers

Review Open access Sep 2026

Performance and Consistency of Large Language Models in Key Labor-Intensive Tasks of Systematic Reviews.

A suite of automated tools for automated batch processing that provide decision rationales and evidence enhances transparency and allows for human verification of AI decisions and provides a suite of automated tools for key SR tasks.

Yi-Ran Liu, Xi-Ling Wang, Zi-Xuan Zhou et al. · 0 citations
#large language models Review Sep 2026

Large language models for literature screening in conceptually complex and interdisciplinary reviews

This study examines the capabilities and limitations of large language models (LLMs) for literature screening in systematic reviews involving conceptually diffuse and interdisciplinary topics. Two screening datasets were constructed. Four LLMs, Claude 4.6, DeepSeek V4 Pro, Gemini 3.1 Pro, and GPT-5.5, were e...

Ni Cheng, Heng Dong, Xuan Han et al. · 0 citations
Review Open access Oct 2026

Title and abstract screening for systematic reviews with Jev, a System One model: comparison with generative large language models

Large language models (LLMs) screen titles and abstracts without review-specific training, but generating screening decisions as text takes processing time and incurs API charges. We evaluated Jev, a non-generative model returning classification probabilities, on 4527 records from two systematic reviews of bipolar diso...

K. Matsui, Y. Takaesu · 0 citations

A Benchmark Framework for Screening Automation in Systematic Reviews

An evaluation framework that accounts for class imbalance is proposed, i.e., the natural prevalence of excluded articles relative to included articles in SRs, and PromptSR, a tool designed to support prompt experimentation, experiment management, and result analysis for LLM-based screening are introduced.

G.Aravind Kumar, Luciano Marchezan, G. Genois et al. · 0 citations
Open access Sep 2026

Diagnostic Accuracy and Alignment With Human Reader Responses Across Five Large Language Models on Complex Clinical Case Challenges

Introduction: Large language models (LLMs) match or exceed physician accuracy on benchmark diagnostic tasks. Whether higher accuracy reflects human-like diagnostic behavior, which bears on how safely clinicians can supervise these systems, is rarely assessed. Objectives: To determine whether diagnostic accuracy and al...

R. Bellocco, L. Soraci, Lorenzo Lo Cicero et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.