Skip to content
Review Open access

LLM-Assisted Reviewer Assignment via Auditable Expertise Matching

2026 · IEEE Access · Vol 14, pp. 127834-127863 · 0 citations · 63 references

TL;DR

Results highlight consistent trade-offs across representations and matchers: two-stage re-ranking improves early-rank performance; keyword-aware Sentence-BERT increases top-3 concentration; and KG edge overlap is competitive on full abstracts, while some graph variants substantially concentrate reviewer workloads.

Abstract

Accurate reviewer assignment is essential for scalable peer review, yet practical systems must balance semantic match quality with transparency, auditability, and manageable reviewer workloads. We study an end-to-end, LLM-assisted reviewer–paper matching pipeline in which large language models are used primarily for affinity construction by extracting and normalizing high-salience topical key phrases from submissions and reviewer publication summaries. We implement an agentic ensemble of heterogeneous producer models to generate candidate keyword sets, followed by a stronger judge agent that selects the best candidate for each input and representation, yielding compact, machine-parseable keyword profiles. Using a dataset of 663 submissions from 85 computer-science venues and 524 reviewer profiles, we evaluate three representations (titles, abstract summaries, and full abstracts) and compare lexical, neural, and graph-based affinity estimators. These include TF–IDF and Jaccard scoring (and their mean), Sentence-BERT similarity, keyword-aware retrieval with two-stage bi-encoder/cross-encoder re-ranking, and graph-based matchers over knowledge graphs and document-local graphs using Node2Vec embeddings with Jaccard-style fusion. We report MRR@3, Precision@3, and MAP@3. In addition, we simulate conference-level reviewer assignment to evaluate workload realism and inequality, summarizing reviewer loads both per-conference (macro) and globally (pooled), and quantifying inequality with the Gini coefficient. Results highlight consistent trade-offs across representations and matchers: two-stage re-ranking improves early-rank performance; keyword-aware Sentence-BERT increases top-3 concentration; and KG edge overlap is competitive on full abstracts, while some graph variants substantially concentrate reviewer workloads.

Read PDF

Similar papers

Conference Aug 2026

XRAG-CJG: A Trust-Aware and Explainable Retrieval Framework for Reviewer Recommendation

Assigning suitable reviewers to academic submissions remains challenging due to the limitations of similarity-based methods and the limited interpretability of modern learning-based approaches. This paper proposes XRAG-CJG, a trust-aware and explainable reviewer recommendation framework that combines semantic retrieval...

Harini Gunawardana, Thushari Silva · 0 citations
#artificial intelligence Review Sep 2026

More Than Mimicking Reviewers: Evaluating LLMs for Pre-Submission Peer Review

Peer-review feedback often arrives too late for authors to make meaningful revisions. We study an author-facing LLM system that moves part of this stress test before submission: it generates a broad pool of atomic concerns and compresses them into a short report. We evaluate agreement with historical reviews and, separ...

Pouya Parsa, Amin Rezaei · 0 citations
Book Open access Aug 2026

SurveyReview: A Reviewer-Aligned Benchmark for Survey Evaluators

This work proposes SurveyReview, a reviewer-aligned, multi-dimensional benchmark and dataset for survey evaluation, and develops a strong baseline evaluator that substantially improves alignment with human reviewers, providing a competitive reference for future research.

Yuheng Zhang, Yuanchun Wang, Fanjin Zhang et al. · 0 citations
Review Open access Aug 2026

LLM aspect prediction: reviewing academic papers from different aspects with Large Language Model

Peer review is a fundamental process in scholarly publishing, wherein reviewers assess and score various aspects of a manuscript (e.g., novelty, clarity, and significance) based on established evaluation criteria. However, this process demands substantial time and effort, and remains inherently susceptible to human bia...

Zi-Hao Hu, F. Fukumoto, Jian He et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.