Skip to content
Review Open access

Evidence-Aware Human-in-the-Loop LLM Review for Requirements-to-Planning Decisions

Sep 2026 · Computers · Vol 15, pp. 608 · 0 citations · 25 references

TL;DR

ReqPlan-Eval is presented, an evidence-aware human-in-the-loop architecture that connects NFR disagreement, weak-word cues, planning-relevant ambiguity, and role-specialized hypotheses to inspectable planning-support records and configurable review routes.

Abstract

Large language models can produce fluent requirements refinements and planning artifacts while still leaving information unresolved for implementation, testing, or planning commitment. This paper presents ReqPlan-Eval, an evidence-aware human-in-the-loop architecture that connects NFR disagreement, weak-word cues, planning-relevant ambiguity, and role-specialized hypotheses to inspectable planning-support records and configurable review routes. The empirical study evaluates the principal mechanisms and role-based routing signals on separate datasets; it does not constitute an end-to-end evaluation of the complete pipeline on a common set of requirements. A held-out 500-requirement NFR diagnostic achieved exact-match accuracy of 0.716, micro F1 of 0.702, and quality accuracy of 0.874. Pattern-aware arbitration increased weak-word specificity from 0.572 to 0.676 and review precision from 0.678 to 0.724, while component/goal gating increased ambiguity specificity from 0.380 to 0.908 and F1 from 0.748 to 0.855; both mechanisms lost recall. In Experiment 3, four role-specialized outputs were compared with a model-seeded reference reviewed and adjudicated by three human reviewers. On the 20 public stories, role-level exact agreement for validation_needed was 0.825, and a two-or-more-vote policy reviewed 12 stories and captured 12 of 14 reference positives without routing any of the six reference negatives. Open-ended planning artifacts showed very low exact normalized-item overlap with the reference for tasks, acceptance criteria, and test ideas, so no semantic-agreement claim is made for those artifacts. Seven external participants provided initial face-validity evidence for selective review and human control. The results support the evaluated component mechanisms and selective review routing under the frozen configurations, but they do not establish end-to-end workflow effectiveness, model invariance, better planning decisions, or industrial effectiveness.

Read PDF

Similar papers

#artificial intelligence Review Sep 2026

Checkpoints Are Not Enough: Trust Calibration in CoSLR, a Human-AI System for Systematic Literature Reviews

CoSLR is presented, a Human-AI collaborative multi-agent system that supports the SLR workflow through a modular three-phase pipeline using large language models and Retrieval-Augmented Generation, and that places explicit, mandatory human checkpoints on the path between generated output and its acceptance.

Aidul Islam, M. Sami, Muhammad Waseem et al. · 0 citations
#natural language process... Preprint Sep 2026

MTAC-IFBench: Benchmarking Instruction-Following in Multi-Turn Agentic Coding

Recently, the rapid development of large language models (LLMs) has reshaped software engineering by enabling autonomous code agents that plan, execute, and utilize external tools iteratively to tackle complex tasks. Beyond achieving functional correctness, these agents must faithfully follow process instructions and c...

Bo-Si Wen, Cun-Xiang Wang, Jia-Yi Gui et al. · 0 citations
#human-computer interacti... Preprint Sep 2026

A Scenario-Knowledge-Driven Pipeline for Just-in-Time Assistance

A scenario-knowledge-driven pipeline is proposed: a single scenario knowledge document, human-authored and version-controlled, configures sensing, constrains LLM reasoning, and shapes a graded intervention proposal that appropriateness of these interventions, the pipeline's restraint on sessions without struggle, and t...

Zhi-Yuan Li, T. Hara, Jun Ota · 0 citations
Conference Open access Sep 2026

PENTESTLLMAGENT: A Task Dependency Graph Planning-Based Multi-Agent Framework for Automated Penetration Testing

PentestLLMAgent is proposed, which integrates a Task Dependency Graph (TDG) for dynamic planning and backtracking; a Hierarchical Multi-Agent Architecture (HMA) with function-calling-based tool invocation, output filtering, and semantic compression, and Executable Knowledge-Guided Command Generation (EKG-CG) for retrie...

Shuo Sheng, Jixin Zhang, Jia Yang et al. · 0 citations
Sep 2026

Release-Aware Prior-Guided Regex Planning for Token-Bounded Log Retrieval in AIOps

Large language models are increasingly being explored for incident triage and root-cause analysis in AIOps, but their practical use in cloud operations is constrained by the volume of logs produced during incident windows. In large distributed systems, a single fault can generate hundreds of thousands of log messages a...

Sharan Babu Paramasivam Murugesan · 0 citations
Review Sep 2026

Human-AI Collaboration for Multi-Line Task Adjustment Using Local Large Language Models and a Digital Twin

Automation systems must adapt to changing tasks, equipment states, and staffing conditions while providing evidence for human review. This study presents a multi-line task-adjustment system integrating a local large language model, a digital twin, and human decision-making. A Propose-Verify-Decide workflow translates o...

Teng-Hsien Ko, Chin-Te Lin · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.