Skip to content
#artificial intelligence Review Open access

Evaluating the use of artificial intelligence (AI) in systematic review abstract screening: a comparative study of freely available AI-aided tools

Sep 2026 · Systematic Reviews · 0 citations

Abstract

Manual abstract screening in systematic reviews is a time-consuming and labour-intensive task. With the rise of artificial intelligence (AI), the number of published articles has grown substantially, adding to the workload of review studies that rely on robust and timely evidence synthesis. At the same time, AI-aided screening tools have been developed to accelerate this process. While previous studies have demonstrated the efficiency of such tools, ongoing technological advances necessitate updated evaluations, particularly for tools that are freely available. In review types such as umbrella reviews, where both the topic area and study design are central to eligibility decisions, the performance of AI-aided tools remains underexplored. We conducted a comparative evaluation of six freely available AI-aided abstract screening tools: Rayyan, RobotAnalyst, PICO Portal, Abstrackr, ASReview, and Colandr using a previously completed umbrella review of interdisciplinary urban planning and public health studies. We assessed (1) early recall performance (i.e., identification of included studies within the first 10% and 25% of screening), (2) feature availability and depth, and (3) user experience. This Study Within a Review (SWAR) was registered in the SWAR repository as SWAR 25. All evaluated tools supported the review process by facilitating screening and offering features such as prioritization and keyword highlighting. However, none identified more than 50% of the previously included studies within the first 25% of screening. Feature analysis and user feedback suggested that Rayyan and PICO Portal achieved the highest feature analysis scores among the evaluated tools for our interdisciplinary umbrella review context, although limitations were noted in duplicate removal and in recognizing the importance of study design in eligibility decisions. Although a growing number of AI-aided abstract screening tools are publicly and freely available, their accuracy, usability, and adaptability to different review designs remain limited. Enhanced support for duplicate detection and integration of study design considerations could improve their utility in umbrella reviews and other complex evidence syntheses. Continued evaluation and user training may support broader adoption across diverse research contexts.

Read PDF

Similar papers

#artificial intelligence Open access May 2023

Evaluating the Performance of Large Language Models on GAOKAO Benchmark

GAOKAO-Bench is introduced, an intuitive benchmark that employs questions from the Chinese GAOKAO examination as test samples, including both subjective and objective questions that contribute a robust evaluation benchmark for future large language models and offers valuable insights into the advantages and limitations...

Xiaotian Zhang, Chun-yan Li, Yi Zong et al. · 216 citations · ⚡17
#artificial intelligence Open access Jul 2024

Gender, Race, and Intersectional Bias in Resume Screening via Language Model Retrieval

This work investigates the possibilities of using LLMs in a resume screening setting via a document retrieval framework that simulates job candidate selection and finds that the MTEs are biased, significantly favoring White-associated names in 85% of cases and female-associated names in only 11.1% of cases.

Kyra Wilson, Aylin Caliskan · 131 citations · ⚡8
#artificial intelligence Review Oct 2025

Ultralytics YOLO Evolution: An Overview of YOLO26, YOLO11, YOLOv8 and YOLOv5 Object Detectors for Computer Vision and Pattern Recognition

This paper presents a comprehensive overview of the Ultralytics YOLO family, emphasizing architectural evolution, benchmarking, deployment, and emerging directions from YOLOv5 through YOLO27, and examines detection, segmentation, depth, classification, pose, oriented detection, tracking, export, quantization, and deplo...

Ranjan Sapkota, Manoj Karkee · 112 citations · ⚡10

The Death of Schema Linking? Text-to-SQL in the Age of Well-Reasoned Language Models

This work revisits schema linking when using the latest generation of large language models (LLMs) and finds empirically that newer models are adept at utilizing relevant schema elements during generation even in the presence of large numbers of irrelevant ones.

Karime Maamari, Fadhil Abubaker, Daniel Jaroslawicz et al. · 109 citations · ⚡19

BadRAG: Identifying Vulnerabilities in Retrieval Augmented Generation of Large Language Models

A novel threat is unveiled in which attackers steer the RAG system's response by injecting malicious passages into its knowledge base, enabling the attacker to steer the response without altering the user input or modifying the RAG weights.

Jiaqi Xue, Meng Zheng, Yebowen Hu et al. · 109 citations · ⚡8

PRISM: Self-Pruning Intrinsic Selection Method for Training-Free Multimodal Data Selection

Empirically, PRISM reduces the end-to-end time for data selection and model tuning to just 30% of conventional pipelines, and achieves this efficiency while simultaneously enhancing performance, surpassing models fine-tuned on the full dataset across eight multimodal and three language understanding benchmarks.

Jinhe Bi, Yifan Wang, Danqi Yan et al. · 73 citations · ⚡4

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.