Aug 2026· IEEE International Requirements Engineering Conference· pp. 470-477· 0 citations· 23 references
Abstract
Large Language Models (LLMs) are increasingly applied to requirements engineering tasks, yet existing benchmarks measure model performance without accounting for how properties of the input text affect extraction outcomes. Among the broad spectrum of requirements engineering activities, this research focuses on requirements extraction, identifying requirement statements in existing documents, a task that can be evaluated against human-annotated ground truth. We argue that measurable text properties can indicate extraction performance. We propose a three-stage research program. Stage 1 defines a metrics toolkit spanning five dimensions (entity coherence, readability, structure, requirements readiness, domain terminology) and validates their discriminative power on a corpus of 376 documents: 19 metrics were used for statistical analysis; all 19 significantly distinguish document types (Welch's ANOVA with Benjamini-Hochberg FDR correction, $p<0.05$), achieving $93.09 \%$ classification accuracy. Stage 2 will assemble a dataset of natural-language documents paired with human-annotated requirement statements, then measure how each document's metric profile correlates with LLM extraction quality. Stage 3 will compare multiple LLMs and prompting strategies to map technique-specific failure boundaries. The envisioned outcome is a decision framework mapping document profiles to expected extraction quality. We report Stage 1 results and the research roadmap.
Whether measurable linguistic properties of prompts can predict LLM performance before inference is investigated, enabling low-cost prompt selection and refinement, validated on binary requirements classification targeting F1, F2, precision, and recall.
Quim Motger, Alessio Miaschi, Xavier Franch et al.· 1 citation
Large language models (LLMs) are increasingly researched as a support for complex requirements engineering (RE) tasks such as requirements elicitation, analysis, and modeling. A fundamental aspect of such tasks is the ability to reliably recognize which parts of natural language texts convey (functional) requirement in...
Jan Böttcher, Klaus Schmid· 2026 IEEE 34th International...· 0 citations
The adoption of large language models (LLMs) in software engineering has enabled the potential to automate complex activities such as requirements analysis. This paper presents an empirical performance analysis of four modern LLMs: GPT-4o, Aya, Gemma and Phi-4 on the task of automated classification of atomic software...
Nourchène Elleuch Ben Ayed, Jaber Jemai, Keletso J. Letsholo et al.· Journal of Information &...· 0 citations
A lightweight quality-assessment protocol is presented for LLM-generated synthetic training data and applied to 13,579 synthetic user reviews generated from GitHub issues across four open-source Android applications, high-lighting the need for hybrid human-AI verification when synthetic data is used in security-critica...
An LLM-based framework is proposed that leverages full-text key-insight extraction to enhance literature classification and implemented a confidence-weighted voting (CWV) mechanism using multiple LLMs to improve robustness.
Zihan Song, Shan Huang, Ngeemasara Thapa et al.· 0 citations
The findings indicate that open-source LLMs can support domain-specific scholarly metadata classification without task-specific fine-tuning, however, their moderate and dimension-dependent performance limits their suitability for fully automated fine-grained metadata enrichment.
M. Haris, Maryam Badar· Frontiers in Research Metric...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.