Skip to content
Review

Beyond Lexical Metrics: Sentence-Embedding Detection of Reviewer Habituation in AI Code Review

Sep 2026 · 0 citations · 20 references
Computer Science

TL;DR

Code review is a key quality checkpoint between AI-generated code and production and reviewer adaptation is detectable in latent distributional structure rather than in classical lexical metrics, and language shifts follow rather than precede changes in approval behavior.

Abstract

Code review is a key quality checkpoint between AI-generated code and production. As AI coding agents submit pull requests at scale, it is unclear whether reviewers reduce scrutiny with repeated exposure and whether review comments reveal this change. We study 11,429 reviews from 400 repeat reviewers over 207 days, paired with 10,104 human-authored inline comments from AIDev. Approval rates rise from 30.5% in reviewers'early periods to 36.6% in late periods (Wilcoxon p = 8.6 x 10^-8; Cohen's d = 0.25). However, four hand-crafted linguistic features - lexical diversity, Shannon entropy, technical specificity, and constructive actionability - show no monotonic decline across exposure deciles (all Spearman absolute rho<= 0.53, p>= 0.11; Bonferroni-corrected Mann-Whitney p>= 0.36). A logistic-regression classifier based on these features reaches F1 = 0.485, below the majority-class baseline. Sentence-embedding structure does carry signal: reviewers'late-period comment centroids shift farther from their early-period centroids than under within-reviewer random permutations (Wilcoxon p<0.001), and a small MLP using three embedding statistics reaches F1 = 0.74 under 5-fold reviewer-stratified cross-validation. Granger analysis shows that approval-rate changes predict later shifts in technical specificity at all tested lags (p<0.001), while the reverse direction is significant at only one of four lags. Reviewer adaptation is therefore detectable in latent distributional structure rather than in classical lexical metrics, and language shifts follow rather than precede changes in approval behavior.

View source

Similar papers

#machine learning Review Sep 2026

Do Reviewers Still Reward Lexical Complexity? A Frozen-Rater Study of Preference Drift in 124K ICLR Reviews

Large language models have collapsed the cost of producing lexically elaborate prose, and whether peer reviewers still reward it is a question about the evaluator, not about the text. When the association between a writing cue and review scores moves across years, the reviewers may have changed, the submissions may hav...

Jia-Bin Zheng · 0 citations
#machine learning Review Sep 2026

When AI Reviews Train AI Reviewers: Scientific-Judgment Collapse and Mitigation

This study introduces TrustReviewer, an open-source LLM-based system for generating peer reviews of AI and machine learning papers and describes a concrete risk of recursive reviewer training and provides practical interventions for preserving judgment diversity and improving recommendation alignment in AI-assisted sci...

Sy-Tuyen Ho, Ming-Hui Liu, Fu-Rong Huang · 0 citations
Open access Sep 2026

Benchmark-Task Heterogeneity in Misinformation-Related and Human–AI Text Classification

Stylometric features and frozen sentence-transformer embeddings are widely used as low-cost inputs for misinformation-related text classification, but heterogeneous benchmark tasks are rarely compared under a common protocol. We evaluated term frequency–inverse document frequency (TF-IDF), a 22-feature stylometric batt...

Grzegorz Świerk, Rafał Olszowski · 0 citations
Review Open access Aug 2026

Large Language Models in Peer Review: Decision Alignment, Review-Text Characteristics, and Human–AI Aggregation at ICLR 2025

Using the ICLR 2025 review process, this study compares 2401 human reviews with 7203 reviews produced in separate, context-isolated API runs using Claude Sonnet 4.5, GPT-5.2 Thinking, and Gemini 3 Pro Preview across decision agreement, review-text characteristics, inter-model consistency, and human–AI aggregation.

Zhi-He Yang, Xiao-Yue Zhou, Hong-Sa Wang et al. · 0 citations
Review Open access Sep 2026

Evaluating large language models as grant reviewers: a comparative study of prompt engineering strategies

Grant application review is resource-intensive and subject to inter-rater variability. Large language models (LLMs) may augment this process, but their reliability in grant evaluation remains unexplored. This exploratory pilot study compared LLM-generated grant reviews to human expert reviews across three prompt engi...

Hants Williams, Jack Evan Lamberg, Eric M. Lamberg · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.