ReqPlan-Eval is presented, an evidence-aware human-in-the-loop architecture that connects NFR disagreement, weak-word cues, planning-relevant ambiguity, and role-specialized hypotheses to inspectable planning-support records and configurable review routes.
Abstract
Large language models can produce fluent requirements refinements and planning artifacts while still leaving information unresolved for implementation, testing, or planning commitment. This paper presents ReqPlan-Eval, an evidence-aware human-in-the-loop architecture that connects NFR disagreement, weak-word cues, planning-relevant ambiguity, and role-specialized hypotheses to inspectable planning-support records and configurable review routes. The empirical study evaluates the principal mechanisms and role-based routing signals on separate datasets; it does not constitute an end-to-end evaluation of the complete pipeline on a common set of requirements. A held-out 500-requirement NFR diagnostic achieved exact-match accuracy of 0.716, micro F1 of 0.702, and quality accuracy of 0.874. Pattern-aware arbitration increased weak-word specificity from 0.572 to 0.676 and review precision from 0.678 to 0.724, while component/goal gating increased ambiguity specificity from 0.380 to 0.908 and F1 from 0.748 to 0.855; both mechanisms lost recall. In Experiment 3, four role-specialized outputs were compared with a model-seeded reference reviewed and adjudicated by three human reviewers. On the 20 public stories, role-level exact agreement for validation_needed was 0.825, and a two-or-more-vote policy reviewed 12 stories and captured 12 of 14 reference positives without routing any of the six reference negatives. Open-ended planning artifacts showed very low exact normalized-item overlap with the reference for tasks, acceptance criteria, and test ideas, so no semantic-agreement claim is made for those artifacts. Seven external participants provided initial face-validity evidence for selective review and human control. The results support the evaluated component mechanisms and selective review routing under the frozen configurations, but they do not establish end-to-end workflow effectiveness, model invariance, better planning decisions, or industrial effectiveness.
CoSLR is presented, a Human-AI collaborative multi-agent system that supports the SLR workflow through a modular three-phase pipeline using large language models and Retrieval-Augmented Generation, and that places explicit, mandatory human checkpoints on the path between generated output and its acceptance.
Aidul Islam, M. Sami, Muhammad Waseem et al.· 0 citations
Recently, the rapid development of large language models (LLMs) has reshaped software engineering by enabling autonomous code agents that plan, execute, and utilize external tools iteratively to tackle complex tasks. Beyond achieving functional correctness, these agents must faithfully follow process instructions and c...
Bo-Si Wen, Cun-Xiang Wang, Jia-Yi Gui et al.· 0 citations
A scenario-knowledge-driven pipeline is proposed: a single scenario knowledge document, human-authored and version-controlled, configures sensing, constrains LLM reasoning, and shapes a graded intervention proposal that appropriateness of these interventions, the pipeline's restraint on sessions without struggle, and t...
PentestLLMAgent is proposed, which integrates a Task Dependency Graph (TDG) for dynamic planning and backtracking; a Hierarchical Multi-Agent Architecture (HMA) with function-calling-based tool invocation, output filtering, and semantic compression, and Executable Knowledge-Guided Command Generation (EKG-CG) for retrie...
Shuo Sheng, Jixin Zhang, Jia Yang et al.· Proceedings of the Thirty-Fi...· 0 citations
Large language models are increasingly being explored for incident triage and root-cause analysis in AIOps, but their practical use in cloud operations is constrained by the volume of logs produced during incident windows. In large distributed systems, a single fault can generate hundreds of thousands of log messages a...
Sharan Babu Paramasivam Murugesan· International Symposium on N...· 0 citations
Automation systems must adapt to changing tasks, equipment states, and staffing conditions while providing evidence for human review. This study presents a multi-line task-adjustment system integrating a local large language model, a digital twin, and human decision-making. A Propose-Verify-Decide workflow translates o...
Teng-Hsien Ko, Chin-Te Lin· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.