Skip to content
Book Open access

SurveyReview: A Reviewer-Aligned Benchmark for Survey Evaluators

Aug 2026 · Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 · 0 citations · 36 references
Computer Science

TL;DR

This work proposes SurveyReview, a reviewer-aligned, multi-dimensional benchmark and dataset for survey evaluation, and develops a strong baseline evaluator that substantially improves alignment with human reviewers, providing a competitive reference for future research.

Abstract

The rapid advancement of large language models has transformed survey writing from a months-long manual effort into an automated process. As generation scales, reliable evaluation becomes the bottleneck, and LLMs are increasingly used as survey evaluators. However, existing approaches largely rely on off-the-shelf LLM-as-a-judge methods without systematic alignment to human reviewers, and there remains a lack of systematic frameworks for quantifying alignment with human reviewers. To address this gap, we propose SurveyReview, a reviewer-aligned, multi-dimensional benchmark and dataset for survey evaluation. We collect and annotate 675 survey papers with 1,630 review reports. We structure authentic peer-review reports by converting free-form comments into four-dimensional scores (Readability, Criticalness, Comprehensiveness, Structure) paired with supporting rationales. We further release standardized train/test splits and an evaluation protocol to measure alignment between automatic evaluators and human reviewers. To validate the benchmark, we develop SurveyAlign, a strong baseline evaluator by fine-tuning Qwen3-32B with LoRA on our annotated data, augmented with external knowledge for knowledge-intensive dimensions. On the test set, SurveyAlign substantially improves reviewer alignment over prompt-based judging with GPT-5.2, reducing average MSE from 2.28 to 1.38 and MAE from 1.15 to 0.69 across all four dimensions. Our contributions are twofold: (1) we establish the first multi-dimensional, reviewer-aligned dataset with a reproducible evaluation framework for survey reviewing; (2) we develop a strong baseline evaluator that substantially improves alignment with human reviewers, providing a competitive reference for future research. Our code and data are available at https://surveyreview.github.io

Read PDF

Similar papers

Review Open access Aug 2026

LLM aspect prediction: reviewing academic papers from different aspects with Large Language Model

Peer review is a fundamental process in scholarly publishing, wherein reviewers assess and score various aspects of a manuscript (e.g., novelty, clarity, and significance) based on established evaluation criteria. However, this process demands substantial time and effort, and remains inherently susceptible to human bia...

Zi-Hao Hu, F. Fukumoto, Jian He et al. · 0 citations
Review Open access 2026

Large Language Model-Based Automated Assessment: A Systematic Review, Taxonomy, and Implications for Personalized Learning

Research on large language model (LLM)-based automated assessment (AA) has expanded rapidly. Nevertheless, the literature remains fragmented across contributions, models, implementation configurations, datasets, and evaluation metrics, complicating efforts to identify approaches suitable for personalized learning. This...

H. D. Septama, A. E. Permanasari, R. Ferdiana · 0 citations
#natural language process... Preprint Sep 2026

When Consistency Does Not Mean Reliability: Evaluating Local LLM Judges Against Human Ratings

The results show that high self-consistency does not necessarily indicate high agreement with human judgments when using local LLMs as automatic judges, and highlight the need to evaluate both consistency and human alignment when using local LLMs as automatic judges.

Aakash Kumar Tiwari · 1 citation
Review Open access 2026

LLM-Assisted Reviewer Assignment via Auditable Expertise Matching

Results highlight consistent trade-offs across representations and matchers: two-stage re-ranking improves early-rank performance; keyword-aware Sentence-BERT increases top-3 concentration; and KG edge overlap is competitive on full abstracts, while some graph variants substantially concentrate reviewer workloads.

Farid Bagheri, Davide Buscaldi, D. Recupero · 0 citations
#natural language process... Preprint Sep 2026

LLJ Cards: Best practices for the Use of LLMs as Judges

In recent years, large language models (LLMs) have emerged as a popular alternative for evaluation. Often referred to as LLMs as judges (LLJs), these systems have been widely adopted by researchers and practitioners across a broad range of measurement tasks, driven by their strong performance, scalability, and cost-eff...

Khaoula Chehbouni, Melina Medjdoub, Florian Carichon et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.