This study introduces TrustReviewer, an open-source LLM-based system for generating peer reviews of AI and machine learning papers and describes a concrete risk of recursive reviewer training and provides practical interventions for preserving judgment diversity and improving recommendation alignment in AI-assisted scientific evaluation.
Abstract
Large language models (LLMs) increasingly participate in scientific evaluation, both as automated reviewers and as assistants to human reviewers. As model-generated reviews enter public data and future training corpora, AI peer review can become recursive: later reviewers learn from judgments produced by earlier models. We study one step of this feedback loop in a controlled setting. Starting from Llama 3.1 8B, we first fine-tune a reviewer on official ICLR reviews from 2018--2023 and then train four successor models on ICLR 2024 data with systematically varied mixtures of official and model-generated reviews. Our study shows that introducing synthetic reviews compresses rating distributions and reduces both same-paper and corpus-level semantic diversity. We call this pattern $\textbf{scientific-judgment collapse}$. To mitigate this failure mode, we introduce $\textbf{TrustReviewer}$, an open-source LLM-based system for generating peer reviews of AI and machine learning papers. TrustReviewer intervenes at two complementary stages. For training-time prevention, we train the core reviewer in a single stage on a curated corpus designed to reduce low-quality and semantically degenerate supervision. For test-time correction, paired activation steering aims to further mitigate residual tendencies toward collapsed judgments without further training or additional expert annotation. Together, these results characterize a concrete risk of recursive reviewer training and provide practical interventions for preserving judgment diversity and improving recommendation alignment in AI-assisted scientific evaluation.
Peer review is a fundamental process in scholarly publishing, wherein reviewers assess and score various aspects of a manuscript (e.g., novelty, clarity, and significance) based on established evaluation criteria. However, this process demands substantial time and effort, and remains inherently susceptible to human bia...
Zi-Hao Hu, F. Fukumoto, Jian He et al.· Scientometrics· 0 citations
Using the ICLR 2025 review process, this study compares 2401 human reviews with 7203 reviews produced in separate, context-isolated API runs using Claude Sonnet 4.5, GPT-5.2 Thinking, and Gemini 3 Pro Preview across decision agreement, review-text characteristics, inter-model consistency, and human–AI aggregation.
Zhi-He Yang, Xiao-Yue Zhou, Hong-Sa Wang et al.· Publications· 0 citations
This survey provides a comprehensive overview of recent advances in LLM-based evaluation, covering techniques, applications, and challenges across domains, with future directions emphasizing standardized protocols, uncertainty estimation, and human–AI collaboration.
M. Nadăş· Artificial Intelligence Revi...· 0 citations
Threshold and judge diagnostics further clarify the limits of interpreting realization as precise anticipation and provides a common task for studying which research ideas a community subsequently pursues and how reliably this outcome can be measured.
Fenghai Li, Zi-Han Tang, Hao-Fei Yu et al.· 1 citation
Grant application review is resource-intensive and subject to inter-rater variability. Large language models (LLMs) may augment this process, but their reliability in grant evaluation remains unexplored. This exploratory pilot study compared LLM-generated grant reviews to human expert reviews across three prompt engi...
Hants Williams, Jack Evan Lamberg, Eric M. Lamberg· Frontiers in Education· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduOct 2, 2026
Martin Trust Center Managing Director Bill Aulet introduces Dear Dreamer, a free platform for middle and high school students who want to learn about entrepreneurship.
Microsoft Research Blog· microsoft.comSep 30, 2026
Extreme space-weather events can damage power systems on Earth and degrade GPS accuracy and satellite operations. A new machine learning system can predict where damage is likely to occur 30-60 minutes before a storm arrives. The post Forecasting space weather risks on power grids appeared first on Microsoft Research.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.