Code review is a key quality checkpoint between AI-generated code and production and reviewer adaptation is detectable in latent distributional structure rather than in classical lexical metrics, and language shifts follow rather than precede changes in approval behavior.
Abstract
Code review is a key quality checkpoint between AI-generated code and production. As AI coding agents submit pull requests at scale, it is unclear whether reviewers reduce scrutiny with repeated exposure and whether review comments reveal this change. We study 11,429 reviews from 400 repeat reviewers over 207 days, paired with 10,104 human-authored inline comments from AIDev. Approval rates rise from 30.5% in reviewers'early periods to 36.6% in late periods (Wilcoxon p = 8.6 x 10^-8; Cohen's d = 0.25). However, four hand-crafted linguistic features - lexical diversity, Shannon entropy, technical specificity, and constructive actionability - show no monotonic decline across exposure deciles (all Spearman absolute rho<= 0.53, p>= 0.11; Bonferroni-corrected Mann-Whitney p>= 0.36). A logistic-regression classifier based on these features reaches F1 = 0.485, below the majority-class baseline. Sentence-embedding structure does carry signal: reviewers'late-period comment centroids shift farther from their early-period centroids than under within-reviewer random permutations (Wilcoxon p<0.001), and a small MLP using three embedding statistics reaches F1 = 0.74 under 5-fold reviewer-stratified cross-validation. Granger analysis shows that approval-rate changes predict later shifts in technical specificity at all tested lags (p<0.001), while the reverse direction is significant at only one of four lags. Reviewer adaptation is therefore detectable in latent distributional structure rather than in classical lexical metrics, and language shifts follow rather than precede changes in approval behavior.
Large language models have collapsed the cost of producing lexically elaborate prose, and whether peer reviewers still reward it is a question about the evaluator, not about the text. When the association between a writing cue and review scores moves across years, the reviewers may have changed, the submissions may hav...
This study introduces TrustReviewer, an open-source LLM-based system for generating peer reviews of AI and machine learning papers and describes a concrete risk of recursive reviewer training and provides practical interventions for preserving judgment diversity and improving recommendation alignment in AI-assisted sci...
Stylometric features and frozen sentence-transformer embeddings are widely used as low-cost inputs for misinformation-related text classification, but heterogeneous benchmark tasks are rarely compared under a common protocol. We evaluated term frequency–inverse document frequency (TF-IDF), a 22-feature stylometric batt...
Grzegorz Świerk, Rafał Olszowski· Information· 0 citations
If a language model can recognize code it wrote, it may favor that code as a judge, and instances of one model monitoring each other could collude, and instances of one model monitoring each other could collude.
Using the ICLR 2025 review process, this study compares 2401 human reviews with 7203 reviews produced in separate, context-isolated API runs using Claude Sonnet 4.5, GPT-5.2 Thinking, and Gemini 3 Pro Preview across decision agreement, review-text characteristics, inter-model consistency, and human–AI aggregation.
Zhi-He Yang, Xiao-Yue Zhou, Hong-Sa Wang et al.· Publications· 0 citations
Grant application review is resource-intensive and subject to inter-rater variability. Large language models (LLMs) may augment this process, but their reliability in grant evaluation remains unexplored. This exploratory pilot study compared LLM-generated grant reviews to human expert reviews across three prompt engi...
Hants Williams, Jack Evan Lamberg, Eric M. Lamberg· Frontiers in Education· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.