Aug 2026· Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2· 0 citations· 36 references
Computer Science
TL;DR
This work proposes SurveyReview, a reviewer-aligned, multi-dimensional benchmark and dataset for survey evaluation, and develops a strong baseline evaluator that substantially improves alignment with human reviewers, providing a competitive reference for future research.
Abstract
The rapid advancement of large language models has transformed survey writing from a months-long manual effort into an automated process. As generation scales, reliable evaluation becomes the bottleneck, and LLMs are increasingly used as survey evaluators. However, existing approaches largely rely on off-the-shelf LLM-as-a-judge methods without systematic alignment to human reviewers, and there remains a lack of systematic frameworks for quantifying alignment with human reviewers. To address this gap, we propose SurveyReview, a reviewer-aligned, multi-dimensional benchmark and dataset for survey evaluation. We collect and annotate 675 survey papers with 1,630 review reports. We structure authentic peer-review reports by converting free-form comments into four-dimensional scores (Readability, Criticalness, Comprehensiveness, Structure) paired with supporting rationales. We further release standardized train/test splits and an evaluation protocol to measure alignment between automatic evaluators and human reviewers. To validate the benchmark, we develop SurveyAlign, a strong baseline evaluator by fine-tuning Qwen3-32B with LoRA on our annotated data, augmented with external knowledge for knowledge-intensive dimensions. On the test set, SurveyAlign substantially improves reviewer alignment over prompt-based judging with GPT-5.2, reducing average MSE from 2.28 to 1.38 and MAE from 1.15 to 0.69 across all four dimensions. Our contributions are twofold: (1) we establish the first multi-dimensional, reviewer-aligned dataset with a reproducible evaluation framework for survey reviewing; (2) we develop a strong baseline evaluator that substantially improves alignment with human reviewers, providing a competitive reference for future research. Our code and data are available at https://surveyreview.github.io
Peer review is a fundamental process in scholarly publishing, wherein reviewers assess and score various aspects of a manuscript (e.g., novelty, clarity, and significance) based on established evaluation criteria. However, this process demands substantial time and effort, and remains inherently susceptible to human bia...
Zi-Hao Hu, F. Fukumoto, Jian He et al.· Scientometrics· 0 citations
Research on large language model (LLM)-based automated assessment (AA) has expanded rapidly. Nevertheless, the literature remains fragmented across contributions, models, implementation configurations, datasets, and evaluation metrics, complicating efforts to identify approaches suitable for personalized learning. This...
H. D. Septama, A. E. Permanasari, R. Ferdiana· IEEE Access· 0 citations
The results show that high self-consistency does not necessarily indicate high agreement with human judgments when using local LLMs as automatic judges, and highlight the need to evaluate both consistency and human alignment when using local LLMs as automatic judges.
Results highlight consistent trade-offs across representations and matchers: two-stage re-ranking improves early-rank performance; keyword-aware Sentence-BERT increases top-3 concentration; and KG edge overlap is competitive on full abstracts, while some graph variants substantially concentrate reviewer workloads.
Farid Bagheri, Davide Buscaldi, D. Recupero· IEEE Access· 0 citations
In recent years, large language models (LLMs) have emerged as a popular alternative for evaluation. Often referred to as LLMs as judges (LLJs), these systems have been widely adopted by researchers and practitioners across a broad range of measurement tasks, driven by their strong performance, scalability, and cost-eff...
Khaoula Chehbouni, Melina Medjdoub, Florian Carichon et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.