Skip to content

PRISMA-LLM: An Empirical Reporting Framework for AI-Assisted Systematic Reviews

Sep 2026 · 0 citations · 19 references
Computer Science

TL;DR

This work analyzes SciLitBench, a corpus of 888 review-automation papers with 14,726 annotations, to characterize changes in methods, review-stage use, evaluation and reported limitations, and introduces PRISMA-LLM, an empirically grounded framework separating implementation disclosure from consequence-sensitive evaluation and limitation reporting.

Abstract

Large language models (LLMs) and AI-enabled software increasingly participate in systematic-review decisions, yet the information needed to audit these workflows is reported inconsistently. We analyze SciLitBench, a corpus of 888 review-automation papers with 14,726 annotations, to characterize changes in methods, review-stage use, evaluation and reported limitations. Automation has shifted toward LLM- and software-facing workflows, including stages that can alter the evidence base. Since 2023, 38.0% of software/product papers reported no evaluation, compared with 9.3% of LLM papers. Reporting coverage increased with LLM workflow complexity, yet 52% of positive-only LLM evaluations still reported an unmet reliability or performance requirement. From these patterns, we introduce PRISMA-LLM, an empirically grounded framework separating implementation disclosure from consequence-sensitive evaluation and limitation reporting.

View source

Similar papers

Sep 2026

Automated LLM-based Classification of Software Requirements

The adoption of large language models (LLMs) in software engineering has enabled the potential to automate complex activities such as requirements analysis. This paper presents an empirical performance analysis of four modern LLMs: GPT-4o, Aya, Gemma and Phi-4 on the task of automated classification of atomic software...

Nourchène Elleuch Ben Ayed, Jaber Jemai, Keletso J. Letsholo et al. · 0 citations

A Benchmark Framework for Screening Automation in Systematic Reviews

An evaluation framework that accounts for class imbalance is proposed, i.e., the natural prevalence of excluded articles relative to included articles in SRs, and PromptSR, a tool designed to support prompt experimentation, experiment management, and result analysis for LLM-based screening are introduced.

G.Aravind Kumar, Luciano Marchezan, G. Genois et al. · 0 citations
Preprint Aug 2026

Challenges and Contributions in Quality of AI-Based Software: A Systematic Mapping Study

A Systematic Mapping Study on the quality of AI-based software identifies six recurring challenge categories, with the most prominent being limitations in existing quality assessment models followed by issues in non-functional requirement management, quality-aware development, and quality assurance.

Maryum Hamdani, Mateen Ahmed Abbasi, Marko Jäntti et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.