Tree-of-Concerns is introduced, a multi-agent framework that deploys specialized skeptic personas, each operating through a category-specific analytical lens, as parallel debate trees to extract unstated limitations from scientific papers.
Abstract
As scientific literature grows and papers increasingly under-report limitations, multi-agent LLMs offer a promising approach to systematically uncover these hidden failure modes. Here, we introduce Tree-of-Concerns, a multi-agent framework that deploys specialized skeptic personas, each operating through a category-specific analytical lens, as parallel debate trees to extract unstated limitations from scientific papers. Each persona conducts structured, evidence-grounded argumentation, while a Panel Review mechanism re-evaluates each surviving claim from all five perspectives to correct category drift and severity miscalibration. Through retrieval-free, single-paper experiments on ToC-Bench, our benchmark of 414 research papers with 1,905 unstated limitations, sourced from reviewer-reported weaknesses and follow-up citation critiques, we demonstrate that ToC improves precision by 79% and coverage by 11% relative to the strongest baseline, surfacing specific, evidence-grounded concerns that support reviewers in systematic evaluation.
Autoresearch agents are reshaping the research ecosystem, but they can also let flawed claims enter the literature at scale. Human advisors catch such issues in drafts through careful, traceable feedback, yet advisor-style assessment requires extensive manual effort and does not scale. To shift automated paper assessme...
Media bias in news articles operates through subtle linguistic cues---loaded language, selective framing, and strategic omission---that resist single-model detection and have traditionally required large annotated corpora for supervised training. We ask whether structured multi-agent deliberation can serve as a princip...
Garvit Joshi, Stavya Dhyani, Jasmine et al.· 0 citations
It is demonstrated that multi-agent consensus can enforce artificial agreement at the expense of true human alignment at the expense of true human alignment, revealing a structural limitation in consensus-style, role-specialized MAD protocols for subjective scoring.
Mi-Ra Song, Chanwoo Kim, Sugyeong Eo et al.· 0 citations
Automated fact-checking systems still fall short of producing explanations that mirror the depth and structure of expert human reasoning. In this work, we propose a multi-agent framework that integrates five specialized linguistic agents covering polarization, linguistic style, argumentation, plausibility, and contextu...
Pedro Henrique de Oliveira Silva, L. Santos, L. Marinho et al.· Proceedings of the 37th ACM...· 0 citations
AI reviewers can now produce many specific criticisms, but more criticism is not necessarily a better review. A review may miss a consequential weakness or retain an allegation that available evidence does not support. These failures require opposite corrections, yet generation-oriented systems and aggregate measures o...
Ze-Xing Zhang, Ji-Chao Li, Tian-Yang Lei et al.· 0 citations
Multi-agent debate, in which several LLMs exchange arguments before producing an answer, is widely assumed to improve answer quality by surfacing genuine disagreement. That disagreement is hard to verify, and no single signal can settle it, so we organize the analysis around four questions: (A) does the debater say it...
Qian Chen· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.