A Task-Contingent Legitimacy framework offering a task-tiered policy approach and testable propositions is formalized in a Task-Contingent Legitimacy framework offering a task-tiered policy approach and testable propositions.
Abstract
The rapid adoption of large language models (LLMs) has prompted extensive debate about their appropriate role in peer review, scholarly publishing’s primary quality-control mechanism. However, AI has not yet been formally approved as a peer-review tool by most academic journals. This study reviews the emerging AI-in-peer-review literature to identify research trends, synthesize empirical evidence across review tasks, and develop a conceptual framework for AI-assisted review. Using a PRISMA-guided Scopus search (176 records identified, 162 included), we combined three-layer content analysis (theme, editorial stance, and AI autonomy) with a synthesis of 18 empirical studies. The literature expanded from 6 records before 2023 to 45 records in the first half of 2026 and remains dominated by commentary and opinion (57%), with the remaining 43% comprising research studies, technical work, and reviews. Editorial perspectives are generally balanced, and authors overwhelmingly favor assistive, human-in-the-loop AI over human-only or full automation. Empirical evidence shows a task-contingent pattern: AI performs well on narrowly defined evaluative tasks (Pearson r > 0.9 in some settings) but less reliably when predicting editorial decisions (accuracy 40–67%; correlations as low as ρ = 0.00). AI legitimacy may depend more on task type than on any governance position, a pattern we formalize in a Task-Contingent Legitimacy framework offering a task-tiered policy approach and testable propositions.
This systematic review investigates the application of artificial intelligence (AI) and machine learning (ML) in journalism and media practice from 2020 to 2026. Following PRISMA guidelines, we analysed 121 peer-reviewed articles from Scopus and Web of Science using a multi-method approach that combined qualitative the...
Gui Jun, Nasrullah Dharejo, M. Alivi· Journalism and Media· 0 citations
It is demonstrated that aggregate quality scores alone can overestimate review quality and argued for multi-dimensional evaluation of AI-generated peer reviews.
Alexander M. Fichtl, Lukas Ellinger, Josefin Kelber et al.· 1 citation
The findings suggest that modern Large Language Models can provide useful and consistent support for scientific peer review, however remaining differences between AI-generated and human-generated evaluations indicate that current systems should be viewed as complementary tools that assist human reviewers rather than re...
Vuk D. Tomić, T. Heyman, E. V. van Nieuwenburg· 0 citations
The results demonstrate clear competency boundaries for GenAI within the peer review, particularly concerning tasks requiring nuanced professional judgment, particularly concerning tasks requiring nuanced professional judgment.
Zhong-Shi Wang, Meng-Yue Gong· Journal of Scholarly Communi...· 0 citations
Peer review is central to quality control in science. However, existing evaluations of AI-assisted peer review mainly focus on the overall quality of generated reviews or the accuracy of final decisions. They therefore provide limited evidence about whether model decisions are supported by sufficient and reliable revie...
Siming Yuan, Xue-Yi Zhang, Wang-Ze Ni et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.