Skip to content
Review

ARISMA: Guidelines for AI- and LLM-Assisted Systematic Reviews, Scoping Reviews, and Mapping Studies

Aug 2026 · 0 citations · 39 references
Computer Science

TL;DR

ARISMA treats AI as an inspected, benchmarked, logged, and reversible assistant rather than an autonomous reviewer, built around one governing principle: every consequential scientific decision must remain human-interpretable, human-auditable, and human-accountable.

Abstract

Systematic reviews, scoping reviews, mapping studies, and related evidence syntheses are increasingly difficult to conduct with fully manual workflows as search volumes, update cycles, and synthesis requirements continue to expand. At the same time, artificial intelligence, machine learning, and large language models are rapidly entering review practice across query formulation, screening, extraction, categorization, appraisal support, and reporting. Yet the empirical evidence remains uneven, task-dependent, and insufficient to justify unconstrained automation. Existing standards such as PRISMA 2020, PRISMA-S, PRISMA-ScR, PRISMA-P, PRESS, and SWiM remain essential, but none provides an end-to-end operational standard for when AI use is methodologically appropriate, how it should be validated, which review decisions must remain human-led, and how AI involvement should be reported so that readers can audit it. This paper proposes ARISMA, an AI Reporting and Integration standard for Systematic Methods and Analysis. ARISMA treats AI as an inspected, benchmarked, logged, and reversible assistant rather than an autonomous reviewer. It is built around one governing principle: every consequential scientific decision must remain human-interpretable, human-auditable, and human-accountable. The paper contributes a lifecycle taxonomy, process guidance, stepwise recommendations across the review pipeline, a governance and provenance model, a tool-support framework, an AI-integrated reporting checklist, and a validation matrix. It also addresses legal, privacy, infrastructure, and sustainability considerations. The framework was iteratively refined through structured expert consultation. The result is a practical and auditable guideline for responsible AI-assisted evidence synthesis.

View source

Similar papers

PRISMA-LLM: An Empirical Reporting Framework for AI-Assisted Systematic Reviews

This work analyzes SciLitBench, a corpus of 888 review-automation papers with 14,726 annotations, to characterize changes in methods, review-stage use, evaluation and reported limitations, and introduces PRISMA-LLM, an empirically grounded framework separating implementation disclosure from consequence-sensitive evalua...

Miguel Zabaleta, Bai-Han Lin · 0 citations
Review Open access Aug 2026

Artificial Intelligence Resources for the Screening of Titles and Abstracts in Systematic Reviews: A Scoping Review

This review aims to identify current evidence concerning AI use during preliminary SLR reference screening and describes characteristics such as the different metrics used for reporting performance and how the different algorithms, pipelines, workflows or web applications are validated.

A. M. Barragán, Sara Elena Ortiz Bonett, Eliana-Isabel Rodríguez-Grande et al. · 0 citations
#artificial intelligence Review Sep 2026

Checkpoints Are Not Enough: Trust Calibration in CoSLR, a Human-AI System for Systematic Literature Reviews

CoSLR is presented, a Human-AI collaborative multi-agent system that supports the SLR workflow through a modular three-phase pipeline using large language models and Retrieval-Augmented Generation, and that places explicit, mandatory human checkpoints on the path between generated output and its acceptance.

Aidul Islam, M. Sami, Muhammad Waseem et al. · 0 citations

A Benchmark Framework for Screening Automation in Systematic Reviews

An evaluation framework that accounts for class imbalance is proposed, i.e., the natural prevalence of excluded articles relative to included articles in SRs, and PromptSR, a tool designed to support prompt experimentation, experiment management, and result analysis for LLM-based screening are introduced.

G.Aravind Kumar, Luciano Marchezan, G. Genois et al. · 0 citations
Review Open access 2026

Large Language Model-Based Automated Assessment: A Systematic Review, Taxonomy, and Implications for Personalized Learning

Research on large language model (LLM)-based automated assessment (AA) has expanded rapidly. Nevertheless, the literature remains fragmented across contributions, models, implementation configurations, datasets, and evaluation metrics, complicating efforts to identify approaches suitable for personalized learning. This...

H. D. Septama, A. E. Permanasari, R. Ferdiana · 0 citations
Review Sep 2026

Generative Artificial Intelligence for Systematic Literature Reviews: A Good Practices Report of an ISPOR Special Task Force.

OBJECTIVES Systematic literature reviews (SLRs) are foundational to evidence-based medicine, including health technology assessment (HTA) and health economics and outcomes research (HEOR). Generative artificial intelligence (GenAI) tools are increasingly used in SLR workflows, yet no good practice guidance exists. This...

R. Fleurence, Riaz Qureshi, Rakesh Aggarwal et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.