Skip to content
Review Open access

Evidence-Aware Human-in-the-Loop LLM Review for Requirements-to-Planning Decisions

Sep 2026 · Computers · 0 citations · 25 references

Abstract

Large language models can produce fluent requirements refinements and planning artifacts while still leaving information unresolved for implementation, testing, or planning commitment. This paper presents ReqPlan-Eval, an evidence-aware human-in-the-loop architecture that connects NFR disagreement, weak-word cues, planning-relevant ambiguity, and role-specialized hypotheses to inspectable planning-support records and configurable review routes. The empirical study evaluates the principal mechanisms and role-based routing signals on separate datasets; it does not constitute an end-to-end evaluation of the complete pipeline on a common set of requirements. A held-out 500-requirement NFR diagnostic achieved exact-match accuracy of 0.716, micro F1 of 0.702, and quality accuracy of 0.874. Pattern-aware arbitration increased weak-word specificity from 0.572 to 0.676 and review precision from 0.678 to 0.724, while component/goal gating increased ambiguity specificity from 0.380 to 0.908 and F1 from 0.748 to 0.855; both mechanisms lost recall. In Experiment 3, four role-specialized outputs were compared with a model-seeded reference reviewed and adjudicated by three human reviewers. On the 20 public stories, role-level exact agreement for validation_needed was 0.825, and a two-or-more-vote policy reviewed 12 stories and captured 12 of 14 reference positives without routing any of the six reference negatives. Open-ended planning artifacts showed very low exact normalized-item overlap with the reference for tasks, acceptance criteria, and test ideas, so no semantic-agreement claim is made for those artifacts. Seven external participants provided initial face-validity evidence for selective review and human control. The results support the evaluated component mechanisms and selective review routing under the frozen configurations, but they do not establish end-to-end workflow effectiveness, model invariance, better planning decisions, or industrial effectiveness.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.