Do Large Language Models Reliably Recognize Functional Requirement Information?
Abstract
Large language models (LLMs) are increasingly researched as a support for complex requirements engineering (RE) tasks such as requirements elicitation, analysis, and modeling. A fundamental aspect of such tasks is the ability to reliably recognize which parts of natural language texts convey (functional) requirement information. This paper investigates this fundamental capability by evaluating whether LLMs can correctly recognize functional requirement information in RE-related text excerpts. To deliberately isolate the recognition task from more complex RE tasks, we introduce an evaluation dataset and filtering task to investigate the ability of LLMs to filter requirement-relevant information from RE-related text excerpts. The proposed evaluation setup operationalizes this filtering task by prompting the models to select and reproduce the text fragments that convey functional requirement information. Verbatim reproduction serves as a controlled evaluation strategy that enables deterministic comparison against a reference solution and avoids semantic drift introduced by paraphrasing. Our experiments show that despite the obvious simplicity of the task, models are far from perfect in addressing this task. Our findings indicate that model size alone is not a reliable predictor of functional requirement recognition ability. Also variables like reasoning vs. non-reasoning do not seem to have a strong impact. Furthermore, the results reveal that current LLMs struggle to consistently distinguish functional requirement information from content that does not convey functional requirements. Such limitations in information recognition may also indicate constraints affecting the achievable performance in more complex RE tasks.