Skip to content
Preprint

Does ISO-Grounded NFR Specification Improve LLM Code Generation? A Comparison of Rich and Structured Interventions against a Natural-Language Baseline

Aug 2026 · 0 citations · 15 references
Computer Science

TL;DR

ISO-grounded enrichment improves static quality proxies and reduces sensitivity to prompt wording, but does not reliably improve functional correctness; for error handling, extended-test pass rate decreases, suggesting tension between defensive coding patterns and exact-output benchmarks.

Abstract

In LLM-based code generation, Non-Functional Requirements (NFRs) are often specified as terse one-line phrases. We ask whether grounding those specifications in ISO/IEC 25010 Quality Model, either as rich natural-language prose (NL-rich) or as structured JSON (Structured), improves code generated on HumanEval/HumanEval-ET compared to a RobuNFR-style one-line baseline (NL-simple). We evaluate four NFRs (performance, error handling, code smell, readability) with ten prompt variations per condition under a fixed model snapshot and paired non-parametric analysis. Primary finding: ISO-grounded enrichment improves static quality proxies (unreadability density falls across all four NFRs (e.g., Performance 0.88 ->0.69 for NL-rich)) and reduces sensitivity to prompt wording, but does not reliably improve functional correctness; for error handling, extended-test pass rate decreases, suggesting tension between defensive coding patterns and exact-output benchmarks. Secondary finding: when ISO content is held constant, NL-rich and Structured differ negligibly in correctness (|delta|<= 0.023), indicating that semantic content matters more than JSON-vs-prose format. Practitioners should invest in standard-grounded NFR content rather than serialization form. A fully traceable replication package is provided.

View source

Similar papers

Conference Aug 2026

When Iterative Prompting Fails: An Empirical Study of Unit Test Generation with Open-Source LLMs

Large language models (LLMs) have shown promise in automated unit test generation, yet the effectiveness of prompt engineering for small, locally-deployed open-source models remains poorly understood. Following growing interest in local LLM deployment to mitigate data exposure risks, this paper presents a controlled em...

M. Tran, Khang Mai · 0 citations
Preprint Aug 2026

NoTB: Oracle-Free Triage of LLM-Generated RTL via Cross-Model Formal Consensus

NoTB is introduced, an oracle-free triage framework that infers correctness from cross-model formal consensus and demonstrates that formal cross-model agreement provides a reliable basis for high-confidence triage without model-dependent oracles.

Elisavet Lydia Alvanaki, Je Yang, Biruk B. Seyoum et al. · 0 citations
#natural language process... Preprint Aug 2026

When Who You Are Can Change the Code You Get: A Study of Persona-Induced Bias in LLM Code Generation

Large Language Models (LLMs) are widely used as programming assistants, yet it remains unclear whether and how user's demographic information impacts the technical quality of generated code. We conduct a large-scale empirical study of persona-induced bias in LLM-based code generation, focusing a proprietary model (Gemi...

Anubhav Gupta, M. Figueiredo, L. Machado et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.