Skip to content

Checkpoints Are Not Enough: Trust Calibration in CoSLR, a Human-AI System for Systematic Literature Reviews

Sep 2026 · 0 citations · 34 references
Computer Science

TL;DR

CoSLR is presented, a Human-AI collaborative multi-agent system that supports the SLR workflow through a modular three-phase pipeline using large language models and Retrieval-Augmented Generation, and that places explicit, mandatory human checkpoints on the path between generated output and its acceptance.

Abstract

Systematic Literature Reviews (SLRs) are essential for evidence-based research but remain time-consuming, requiring researchers to manage large volumes of publications across planning, screening, analysis, and reporting. Large language models (LLMs) can now produce fluent, well-structured review text, which makes it difficult to distinguish synthesis that was verified by a researcher from synthesis that merely appears authoritative. This raises the risk that unverified AI-generated synthesis enters the scholarly record carrying the credibility of a systematic review. We present CoSLR, a Human-AI collaborative multi-agent system that supports the SLR workflow through a modular three-phase pipeline using large language models and Retrieval-Augmented Generation (RAG), and that places explicit, mandatory human checkpoints on the path between generated output and its acceptance. In a survey-based study with 63 participants, the system was received positively: 27 of 63 participants (42.9 percent) rated its usability highly, indicating that the mandatory checkpoints did not come at the cost of a workable interface. However, a checkpoint safeguards the review only if researchers use it to verify: 22 of 63 participants (34.9 percent) reported that they would trust AI-generated summaries and reports without additional human checking after only a short interaction with the system. These findings indicate that Human-AI collaboration can support literature review work, but that the effectiveness of human oversight depends on whether users are willing to exercise it. This is a calibration problem that interface design must address directly, not assume.

View source

Similar papers

Review Open access Sep 2026

Evidence-Aware Human-in-the-Loop LLM Review for Requirements-to-Planning Decisions

ReqPlan-Eval is presented, an evidence-aware human-in-the-loop architecture that connects NFR disagreement, weak-word cues, planning-relevant ambiguity, and role-specialized hypotheses to inspectable planning-support records and configurable review routes.

Hamad I. Alsawalqah, Ahmad Abadleh, Shrouq Ibrahim et al. · 0 citations

PRISMA-LLM: An Empirical Reporting Framework for AI-Assisted Systematic Reviews

This work analyzes SciLitBench, a corpus of 888 review-automation papers with 14,726 annotations, to characterize changes in methods, review-stage use, evaluation and reported limitations, and introduces PRISMA-LLM, an empirically grounded framework separating implementation disclosure from consequence-sensitive evalua...

Miguel Zabaleta, Bai-Han Lin · 0 citations
#artificial intelligence Preprint Sep 2026

Research-Native by Construction: Minimal Nodes, Re-verifiable Workflows, and Compounding Memory for Long-Horizon Scientific Agents

The design rests on one claim: most of the credibility of machine-made research can be moved from asking the model to behave to making the non-compliant state unrepresentable, and the system description is a system description written under one rule.

Ding Wang, Yu Liu, Bing Cui et al. · 0 citations
Conference Open access Sep 2026

A Review on Test-Time Scaling for Agentic Large Language Models

A novel RAIE taxonomy along four scaling dimensions is proposed, which optimizes the entire thought process through search algorithms and self-verification, and introduces a task-oriented guideline for choosing the best TTS strategy.

Jia-Yu An, Zheng Chen, Yong-Cheng Jing et al. · 0 citations
Conference 2026

When Verification Hurts: The Cost of Overriding Abstention in Two-Stage Web Agents

This study cautions against transplanting verification into grounding pipelines and identifies calibrated abstention as a property worth preserving and proposes an abstention-aware verifier that intervenes only under sufficient candidate coverage and confidence.

Duchen Li · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.