Large language models are increasingly used for research quality evaluation, with prior work exploring their scoring accuracy and the plausibility of review rationales exploring their scoring accuracy and the plausibility of review rationales.
Abstract
Large language models (LLMs) are increasingly used for research quality evaluation, with prior work exploring their scoring accuracy and the plausibility of review rationales. However, less is known about whether external score guidance changes the evidence presented in the generated review as well as the final score. This study uses 98 Allied Health Professions research outputs submitted for internal REF-style assessment, with specialist human review reports and adjudicated 1-4 reference scores. No-guidance baseline reviews are compared with oracle-guided reviews, where the supplied score is set to the rounded human reference score; extracted evaluation points are used to compare human and LLM evidence use. Using this design, oracle guidance improves scoring accuracy, with score-following checks showing that models do not simply copy the supplied score. Corrected score mismatches are associated with changes in the generated review frame, showing that the score signal can steer review rationales. This effect is direction-dependent: LLM reviews cover human strength or upgrade points more reliably than human weakness or downgrade points, with the weakest alignment for expert downgrade evidence. The results show that score-guided review generation can be evaluated at the level of review evidence, as well as the final score.
Peer-review evaluation is increasingly being automated with LLM-as-a-judge metrics, but this creates a measurement risk. A review may receive a high score because it is fluent, organized, and polished, rather than because it provides a strong evaluation of the paper. This risk is especially important in AI-assisted rev...
Shakiba Amirshahi, Sajad Ebrahimi, Hai-Son Le et al.· 0 citations
Introduction
Recent studies showed poor performance of large language models (LLM) for assessing risk of bias (RoB) with the RoB2 tool. This is in line with the low reliability that humans have in assessing RoB. However, the use of an implementation document (ID) prepared by expert reviewers – i.e., a standardised doc...
C. Del Giovane, M. Marques da Cruz, B. Sousa-Pinto et al.· Epidemiology Biostatistics a...· 0 citations
The findings suggest that modern Large Language Models can provide useful and consistent support for scientific peer review, however remaining differences between AI-generated and human-generated evaluations indicate that current systems should be viewed as complementary tools that assist human reviewers rather than re...
Vuk D. Tomić, T. Heyman, E. V. van Nieuwenburg· 0 citations
BackgroundCochrane Risk of Bias 2 (RoB 2) assessment is methodologically demanding and resource intensive. Large language models may assist this process but can produce variable judgments.
AimTo develop a rule-constrained LLM system for RoB 2 assessment and evaluate its temporal stability and agreement with human revi...
J. T. Joseph, R. Vishwanath, J. V. Sisirkumar· medRxiv· 0 citations
Large language models (LLMs) are increasingly used to generate scientific reviews, yet existing evaluations rarely examine whether different providers align with both conference decisions and human reviewing priorities within the same controlled setting. We compare reviews from OpenAI GPT-5.4, Google Gemini 3.1 Pro Pre...
Abraham Camelo-Guerrero, J. Diaz-Rodriguez· 0 citations
It is demonstrated that aggregate quality scores alone can overestimate review quality and argued for multi-dimensional evaluation of AI-generated peer reviews.
Alexander M. Fichtl, Lukas Ellinger, Josefin Kelber et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.