Aug 2026· 2026 IEEE 34th International Requirements Engineering Conference Workshops (REW)· pp. 5-14· 0 citations· 29 references
Abstract
Behavior-Driven Development (BDD) scenarios coexist with large volumes of semi-structured records (e.g., documentation, issues, informal feature descriptions), and keeping the two synchronized is laborintensive. We present a study of bidirectional generation 11https://github.com/Artin-Biniek/Submission-Project between records and BDD scenarios, comparing finetuning of CodeT5+ encoder–decoder models and the Qwen2.5-Coder-0.5B decoder-only model against retrievalaugmented generation (RAG) with few-shot prompting on Meta-Llama-3.1-8B, CodeT5+, and DeepSeek-Coder-33B. On a curated dataset of 2,100 aligned pairs, finetuning attains best BLEU/F1 of 0.9394/0.9549 and Exact Match of 0.8119 for record generation, while RAG never exceeds BLEU 0.35 and yields Exact Match 0.00. Twoannotator human evaluation, two LLM judges, and twoproportion Z-tests confirm that fine-tuning significantly outperforms RAG $(p<0.001)$ for high-fidelity BDD generation.
Large language models used for code editing can be trained and deployed in at least two output regimes: direct generation, where the model emits the entire modified file in one shot, and iterative diff-based generation ("steps"), where the model emits a sequence of localized search/replace edits applied one at a time u...
Large language models (LLMs) excel at code-centric and algorithmic tasks, but their software design and modeling capabilities remain limited. Unlike abundant code that enables strong generation performance, the scarcity of high-quality design corpora in the training data restricts their software modeling comprehension....
Xiao He, Jing-Wei Shen, Ru Chen· Proceedings of the ACM/IEEE...· 0 citations
A reproducible, human-validated evaluation framework applied to 13 strategies—four architectural families crossed with four reasoning variants crossed with four reasoning variants—across three SLMs spanning 3B–14B parameters, plus targeted ablations.
Balaji Venktesh, Amsaprabhaa M, G. Sundaram· International Conference on...· 0 citations
SpecMine lets the community study, for the first time, how software is specified in the age of AI agents through two censuses: a broad census of spec.md files and a census-wide index of typed references.
Shyam Agarwal, Anmol Singhal, Travis D. Breaux et al.· 0 citations
SchemaGUI, a template-based benchmark for controllable GUI generation evaluation, synthesizing paired natural language instructions and deterministic function-call references from parameterized interface schemas can generate thousands of deterministically annotated tasks in seconds without human labeling.
Jiarui Dong, Yin Cai, Zhouhong Gu et al.· 0 citations
XREPOTEST is introduced, a multilingual repository-level benchmark for unit test generation spanning five underexplored languages: Rust, Go, Julia, PHP, and Ruby, and Invocation Rate is proposed to assess whether generated tests meaningfully exercise the intended functionality.
L. Dung, Dong Cao Van, Nam Le Hai et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.