Skip to content
Review

Metag: A dataset to build agentic meta-reviewing capabilities

Aug 2026 · 0 citations · 28 references
Computer Science

TL;DR

Metag is a dataset to accelerate the development of meta-reviewing agents, specifically to identify changes made to scientific articles during the review-rebuttal process and will enable building methods to empower meta reviewers to quickly identify whether authors have addressed reviewer statements and where in the paper those changes have been made, resulting in additional transparency and traceability throughout peer review.

Abstract

AI tools increasingly support tasks across the scientific research cycle, from experiment design and manuscript preparation to peer review. At the same time, the continuing growth in conference submissions has increased the burden on meta-reviewers, who must synthesize reviewer feedback, author rebuttals, and manuscript revisions. To address this concern, this paper introduces Metag, a dataset to accelerate the development of meta-reviewing agents, specifically to identify changes made to scientific articles during the review-rebuttal process. Each instance contains a reviewer concern, the author's proposed resolution, and the manuscript diffs implementing the stated change. Metag is collected by obtaining manuscript versions from before the review deadline and after acceptance, computing differences between the two documents, and asking human annotators to align these differences with action items from OpenReview discussions. The resulting dataset consists of 349 high-quality action items tied to paper differences and will enable building methods to empower meta reviewers to quickly identify whether authors have addressed reviewer statements and where in the paper those changes have been made, resulting in additional transparency and traceability throughout peer review. The dataset is publicly available at https://github.com/microsoft/Metag-dataset.

View source

Similar papers

Review Open access Oct 2026

An Agentic AI-Aided Review of Large Language Model Applications in Literature Reviews: A Seven-Layer Architecture

The growing volume of scientific output and the pace of AI advancement create a dual challenge: traditional systematic reviews take twelve to eighteen months to complete, risking obsolescence before publication, while the technology needed to accelerate them is itself advancing faster than it can be reviewed. This work...

E. A. Merchán-Cruz, Ioseb Gabelaia, Shwe Soe et al. · 0 citations

When Evidence Conflicts: Reliability-aware Meta-review Generation

Generating coherent meta-reviews from multiple peer reviews is challenging when reviewer evidence conflicts and varies in reliability. Existing approaches typically formulate meta-review generation as a multi-document summarization task and aggregate reviewer feedback uniformly, making it difficult to determine which o...

Xin-Zhe Wang, Fei Tao, Jiang Xie et al. · 0 citations
Review Open access 2026

LLM-Assisted Reviewer Assignment via Auditable Expertise Matching

Results highlight consistent trade-offs across representations and matchers: two-stage re-ranking improves early-rank performance; keyword-aware Sentence-BERT increases top-3 concentration; and KG edge overlap is competitive on full abstracts, while some graph variants substantially concentrate reviewer workloads.

Farid Bagheri, Davide Buscaldi, D. Recupero · 0 citations
Review Open access Sep 2026

Artificial Intelligence in Peer Review: A Bibliometric-Guided Thematic Review and a Task-Contingent Legitimacy Framework

A Task-Contingent Legitimacy framework offering a task-tiered policy approach and testable propositions is formalized in a Task-Contingent Legitimacy framework offering a task-tiered policy approach and testable propositions.

Eungi Kim, Vaishali Singh · 0 citations
#artificial intelligence Review Sep 2026

Checkpoints Are Not Enough: Trust Calibration in CoSLR, a Human-AI System for Systematic Literature Reviews

CoSLR is presented, a Human-AI collaborative multi-agent system that supports the SLR workflow through a modular three-phase pipeline using large language models and Retrieval-Augmented Generation, and that places explicit, mandatory human checkpoints on the path between generated output and its acceptance.

Aidul Islam, M. Sami, Muhammad Waseem et al. · 0 citations
Review Sep 2026

Beyond Human-Likeness: Mapping the Scientific Critique Profiles of LLMs and Human Reviewers

Large language models (LLMs) are increasingly discussed as tools for peer review, but their value is often assessed through human-likeness, perceived usefulness, or textual overlap with reviewer comments. This study shifts attention from whether LLMs resemble human reviewers to what functions of scientific critique the...

YunHong Yang, Mike Thelwall, Guo-Xiu He · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.