Skip to content
Conference

Generating Behavior-Driven Development Artifacts

Aug 2026 · 2026 IEEE 34th International Requirements Engineering Conference Workshops (REW) · pp. 5-14 · 0 citations · 29 references

Abstract

Behavior-Driven Development (BDD) scenarios coexist with large volumes of semi-structured records (e.g., documentation, issues, informal feature descriptions), and keeping the two synchronized is laborintensive. We present a study of bidirectional generation 11https://github.com/Artin-Biniek/Submission-Project between records and BDD scenarios, comparing finetuning of CodeT5+ encoder–decoder models and the Qwen2.5-Coder-0.5B decoder-only model against retrievalaugmented generation (RAG) with few-shot prompting on Meta-Llama-3.1-8B, CodeT5+, and DeepSeek-Coder-33B. On a curated dataset of 2,100 aligned pairs, finetuning attains best BLEU/F1 of 0.9394/0.9549 and Exact Match of 0.8119 for record generation, while RAG never exceeds BLEU 0.35 and yields Exact Match 0.00. Twoannotator human evaluation, two LLM judges, and twoproportion Z-tests confirm that fine-tuning significantly outperforms RAG $(p<0.001)$ for high-fidelity BDD generation.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Diffs vs. Whole Files: An Empirical Comparison of Iterative Edit-Based and Direct Generation for Flutter/Dart Code Models

Large language models used for code editing can be trained and deployed in at least two output regimes: direct generation, where the model emits the entire modified file in one shot, and iterative diff-based generation ("steps"), where the model emits a sequence of localized search/replace edits applied one at a time u...

Andrej Andrejev · 0 citations
Book Open access Oct 2026

LLM-Enhanced Stochastic Generation of Class Diagram Datasets

Large language models (LLMs) excel at code-centric and algorithmic tasks, but their software design and modeling capabilities remain limited. Unlike abundant code that enables strong generation performance, the scarcity of high-quality design corpora in the training data restricts their software modeling comprehension....

Xiao He, Jing-Wei Shen, Ru Chen · 0 citations

Capacity vs. architecture: an evaluation of SLMs for automated docstring generation

A reproducible, human-validated evaluation framework applied to 13 strategies—four architectural families crossed with four reasoning variants crossed with four reasoning variants—across three SLMs spanning 3B–14B parameters, plus targeted ablations.

Balaji Venktesh, Amsaprabhaa M, G. Sundaram · 0 citations
#small language model Preprint Aug 2026

SchemaGUI: A Schema-Driven Benchmark for Controllable GUI Generation Evaluation

SchemaGUI, a template-based benchmark for controllable GUI generation evaluation, synthesizing paired natural language instructions and deterministic function-call references from parameterized interface schemas can generate thousands of deterministically annotated tasks in seconds without human labeling.

Jiarui Dong, Yin Cai, Zhouhong Gu et al. · 0 citations
#software testing Preprint Aug 2026

XREPOTEST: Benchmarking Multilingual Repository-Level Unit Test Generation for Large Language Models

XREPOTEST is introduced, a multilingual repository-level benchmark for unit test generation spanning five underexplored languages: Rust, Go, Julia, PHP, and Ruby, and Invocation Rate is proposed to assess whether generated tests meaningfully exercise the intended functionality.

L. Dung, Dong Cao Van, Nam Le Hai et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.