This work analyzes SciLitBench, a corpus of 888 review-automation papers with 14,726 annotations, to characterize changes in methods, review-stage use, evaluation and reported limitations, and introduces PRISMA-LLM, an empirically grounded framework separating implementation disclosure from consequence-sensitive evaluation and limitation reporting.
Abstract
Large language models (LLMs) and AI-enabled software increasingly participate in systematic-review decisions, yet the information needed to audit these workflows is reported inconsistently. We analyze SciLitBench, a corpus of 888 review-automation papers with 14,726 annotations, to characterize changes in methods, review-stage use, evaluation and reported limitations. Automation has shifted toward LLM- and software-facing workflows, including stages that can alter the evidence base. Since 2023, 38.0% of software/product papers reported no evaluation, compared with 9.3% of LLM papers. Reporting coverage increased with LLM workflow complexity, yet 52% of positive-only LLM evaluations still reported an unmet reliability or performance requirement. From these patterns, we introduce PRISMA-LLM, an empirically grounded framework separating implementation disclosure from consequence-sensitive evaluation and limitation reporting.
ARISMA treats AI as an inspected, benchmarked, logged, and reversible assistant rather than an autonomous reviewer, built around one governing principle: every consequential scientific decision must remain human-interpretable, human-auditable, and human-accountable.
The adoption of large language models (LLMs) in software engineering has enabled the potential to automate complex activities such as requirements analysis. This paper presents an empirical performance analysis of four modern LLMs: GPT-4o, Aya, Gemma and Phi-4 on the task of automated classification of atomic software...
Nourchène Elleuch Ben Ayed, Jaber Jemai, Keletso J. Letsholo et al.· Journal of Information &...· 0 citations
SciLitBench identifies a practical boundary between high-recall screening and evidence-complete extraction and provides a reproducible resource for evaluating LLM-assisted evidence synthesis.
An evaluation framework that accounts for class imbalance is proposed, i.e., the natural prevalence of excluded articles relative to included articles in SRs, and PromptSR, a tool designed to support prompt experimentation, experiment management, and result analysis for LLM-based screening are introduced.
G.Aravind Kumar, Luciano Marchezan, G. Genois et al.· 0 citations
AI-assisted development tools enable software engineers to generate implementations at substantially higher speed and volume than in traditional workflows, yet relatively little is known about how existing guardrails evolve in response.
A Systematic Mapping Study on the quality of AI-based software identifies six recurring challenge categories, with the most prominent being limitations in existing quality assessment models followed by issues in non-functional requirement management, quality-aware development, and quality assurance.
Maryum Hamdani, Mateen Ahmed Abbasi, Marko Jäntti et al.· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduSep 24, 2026
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.