Sep 2026· Journal of Evaluation In Clinical Practice· Vol 32 6, pp.
e70596
· 0 citations· 24 references
Medicine
TL;DR
A suite of automated tools for automated batch processing that provide decision rationales and evidence enhances transparency and allows for human verification of AI decisions and provides a suite of automated tools for key SR tasks.
Abstract
Objective
To evaluate the performance and consistency of Large Language Models (LLMs) in core systematic review (SR) tasks and to introduce open-source tools for automated batch processing that provide decision rationales.
Methods
We assessed GPT-4o, Kimi-K2, DeepSeek-V3, and DeepSeek-R1 on five SR tasks: title/abstract screening (3550 records), full-text screening (233 texts), data extraction (112 RCTs), Risk of Bias (ROB) assessment (112 RCTs), and AMSTAR-2 assessment (20 SRs). Each model was evaluated twice to measure consistency. All outputs required supporting rationales and verbatim evidence.
Results
LLMs demonstrated proficiency across tasks, with generally high intra-model but lower inter-model consistency. In screening, models showed lower precision (0.27-0.40) but high recall (0.83-0.91) and specificity (0.83-0.91). DeepSeek-R1 and DeepSeek-V3 excelled in title/abstract and full-text screening, respectively. Data extraction accuracy was similar across models (0.78-0.82). Kimi-K2 achieved the highest ROB F1 score (0.71). AMSTAR-2 assessments were generally acceptable.
Discussion
While effective, LLMs showed variable performance across SR tasks. The mandatory output of rationales and evidence enhances transparency and allows for human verification of AI decisions.
Conclusion
We provide a suite of automated tools for key SR tasks. By leveraging these tools to validate model outputs rather than starting manually, reviewers can significantly improve workflow efficiency while maintaining methodological rigour.
This study examines the capabilities and limitations of large language models (LLMs) for literature screening in systematic reviews involving conceptually diffuse and interdisciplinary topics.
Two screening datasets were constructed. Four LLMs, Claude 4.6, DeepSeek V4 Pro, Gemini 3.1 Pro, and GPT-5.5, were e...
Ni Cheng, Heng Dong, Xuan Han et al.· Aslib Journal of Information...· 0 citations
Large language model editorial re--ations were prompt sensitive and showed fair agreement across models despite critique statements that were largely grounded in manuscript text, supporting assistive use with human oversight.
S. M. Erturk, Mustafa Durmaz· Academic Radiology· 2 citations
SciLitBench identifies a practical boundary between high-recall screening and evidence-complete extraction and provides a reproducible resource for evaluating LLM-assisted evidence synthesis.
Research on large language model (LLM)-based automated assessment (AA) has expanded rapidly. Nevertheless, the literature remains fragmented across contributions, models, implementation configurations, datasets, and evaluation metrics, complicating efforts to identify approaches suitable for personalized learning. This...
H. D. Septama, A. E. Permanasari, R. Ferdiana· IEEE Access· 0 citations
It is demonstrated that modern LLMs with reasoning capabilities can effectively support real-time clinical evidence synthesis without task-specific fine-tuning, offering a pathway toward scalable automated systems for clinical trial interpretation that could substantially reduce the evidence-to-practice gap in medical...
Shobanapriyan Chandrasegaran, Amal Htait· International Conference on...· 0 citations
Compared human and LLM title-and-abstract screening workflows in a preregistered study embedded in a conceptually complex scoping review, finding that large language models are better suited to validated, auditable, human-supervised workflows than autonomous exclusion.
Nikol Figalová, Lynn Huestegge, Anne Böckler-Raettig· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.