Guiding LLM Peer Reviewers: The Impact of Score Anchors on Review Evidence and Accuracy
Large language models are increasingly used for research quality evaluation, with prior work exploring their scoring accuracy and the plausibility of review rationales exploring their scoring accuracy and the plausibility of review rationales.