Jul 2026· Journal of Information Systems Engineering and Business Intelligence· Vol 12, pp. 443-457· 0 citations· 25 references
TL;DR
An automated natural language processing (NLP)-based framework for requirement elicitation that integrates heterogeneous online sources that transforms heterogeneous textual data into actionable requirement artifacts, providing a scalable and practical solution for early-stage software development.
Abstract
Background: Data-driven requirement elicitation has been increasingly used in modern software engineering due to the growing availability of online user-generated textual data. However, existing approaches mostly rely on single-source data, which are limited in handling the diversity of characteristics found in online textual sources.
Objective: This study proposes an automated natural language processing (NLP)-based framework for requirement elicitation that integrates heterogeneous online sources, such as app reviews, online news, and tweets, for process innovation in the requirements engineering phase.
Methods: The proposed framework combines rule-based and AI-based extraction methods, semantic clustering, and diagram generation. Data were collected from application reviews, Twitter/X, and online news across six domains. The framework was evaluated using expert-annotated ground truth to measure extraction performance and expert-based assessments to examine clustering quality and artifact usefulness.
Result: AI-based extraction outperforms rule-based methods for requirements extraction, achieving F1 scores of 0.92, 0.80, and 0.67 on app reviews, 0.80 on Twitter, and 0.67 on online news. In the expert evaluation, the proposed system demonstrates high topic coherence, reduces elicitation time, and helps identify potential system requirements that may be overlooked in manual processes.
Conclusion: Multisource integration enhances the completeness and contextual richness of automated requirement elicitation. The proposed framework effectively transforms heterogeneous textual data into actionable requirement artifacts, providing a scalable and practical solution for early-stage software development.
Keywords: Requirement Elicitation, Natural Language Processing, Process Innovation, Multisource Data
A multidimensional taxonomy is proposed to classify grounding techniques by architectural paradigm, data source, underlying mechanism, and intervention phase, indicating they are transitioning from an optional improvement to a highly prioritized architectural component for LLM-based recommender systems.
Andrés Felipe Solis Pino, N. Duque-Méndez, Pablo H. Ruiz et al.· Applied Sciences· 0 citations
Large Language Models (LLMs) are increasingly applied to requirements engineering tasks, yet existing benchmarks measure model performance without accounting for how properties of the input text affect extraction outcomes. Among the broad spectrum of requirements engineering activities, this research focuses on require...
Konstantin Valeev, Fabiano Dalpiaz, F. Aydemir· IEEE International Requireme...· 0 citations
A systematic mapping study of 74 peer-reviewed primary studies on AI-based automated requirements elicitation published between 2021 and 2025, identified from five databases following PRISMA 2020 and classified by AI technique, textual source, elicitation activity, and application domain gives researchers and practitio...
Safaa Eltahier, S. Al-Ghuribi, Mawal A. Mohammed et al.· Information· 0 citations
An LLM-based framework is proposed that leverages full-text key-insight extraction to enhance literature classification and implemented a confidence-weighted voting (CWV) mechanism using multiple LLMs to improve robustness.
Zihan Song, Shan Huang, Ngeemasara Thapa et al.· 0 citations
OmniExtract is an automatic data extraction tool with user-friendly configuration files which can adapt to various data extraction tasks and it can support a comprehensive data extraction including text and tables.
The findings show that AI- and NLP-based methods have significantly improved the automation, retrieval, interpretation, and structuring of business documents, and large language models (LLMs), particularly when combined with prompt engineering, retrieval-augmented generation, knowledge graphs, and agent-based architect...
Naif N. Alotaibi, Morteza Saberi, M. Bandara et al.· Analytics· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.