Skip to content

Bridging Behavior and Implementation: Automated Java Glue Code Generation for Behavior-Driven Development

Jul 2026 · arXiv.org · Vol abs/2607.19703 · 0 citations · 51 references
Computer Science

TL;DR

Results show that behavior interpretation and project-aware context retrieval both contribute substantially to generation quality, and demonstrate that LLMs can effectively connect natural-language behavior specifications with project code and support specification-driven software development.

Abstract

Behavior-Driven Development (BDD) helps technical and non-technical stakeholders share a common understanding of software requirements through natural-language scenarios. Glue code makes these scenarios executable by mapping each step to the corresponding project code. However, developing and maintaining glue code requires knowledge of both the intended behavior and the underlying codebase, making it a labor-intensive part of BDD as requirements evolve. Although large language models (LLMs) have shown strong code generation capabilities, their use for automated glue code generation remains unexplored. This task requires reasoning over underspecified behavior, related BDD artifacts, and large project codebases. We present AutoGlue, a hierarchical multi-agent framework for automated Java glue code generation. AutoGlue follows a behavior-first workflow that separates behavior interpretation, context retrieval, and code generation. A Behavior Interpreter derives the intent of a step from its scenario context, while a Developer agent retrieves relevant BDD artifacts and project code before generating the final glue code. We evaluate AutoGlue on 1,307 steps from eight open-source Java projects. Compared with few-shot prompting, AutoGlue improves API F1 by 58.7% and CodeBLEU by 43.7%. It produces directly usable glue code for 46.1% of the evaluated steps, while most partially correct outputs require only minor revisions, such as adding missing actions or refining parameters. Ablation results show that behavior interpretation and project-aware context retrieval both contribute substantially to generation quality. These findings demonstrate that LLMs can effectively connect natural-language behavior specifications with project code and support specification-driven software development.

View source

Similar papers

Jul 2026

TraceDev: A Traceability-Driven Multi-agent Framework for Requirement-to-Code Development

This work proposes TraceDev, a multi-agent framework for automated software development grounded in use cases that contain multiple functional points and complex semantics, and demonstrates the effectiveness of TraceDev in repository-level code generation from requirements.

Mingyu Chen, Ya-Kun Zhang, Zihao Xie et al. · 0 citations
#software testing Preprint Aug 2026

Repo0: Design-Driven Zero-to-All Code Generation

Repo0 is presented, a continuous structural evolution framework for zero-to-all code generation that maintains an explicit architectural state instantiated as a Dual-Directed-Acyclic-Graph (Dual-DAG), consisting of a requirement-level DAG, a component-level DAG, and their alignment relation.

Si-Lin Chen, Haoyi Teng, Xiao-Dong Gu et al. · 2 citations
#artificial intelligence Preprint Sep 2026

WiseSpec: Requirements-Driven Agents for Code Generation

W WiseSpec is proposed, a novel requirements-driven agent framework for repository-level code generation that automatically constructs structured and information-rich requirements, assesses their quality through execution-based evaluation, and iteratively refines them to better guide code generation.

Zhao Tian · 0 citations
Open access 2026

Models, Prompts, and Code! A Semi-Formal State Machine Language for Multi-Paradigmatic Software Development

Compared to both classical UML tooling and fully LLM-based generation, the approach offers stronger determinism, better traceability, lower cognitive modeling effort, and reduced computational cost, while retaining the flexibility to express complex action behavior in natural language where formal specification would b...

Oliver Engling, Felix Schwägerl, Thomas Buchmann · 0 citations

AI-Generated Code Is Not Reproducible (Yet): An Empirical Study of Execution Reliability in LLM-Based Coding Agents

An empirical study of whether complete software artifacts generated by LLM coding agents can be executed in a clean environment using only the code, dependency specifications, and instructions the agent provides suggests that coding-agent evaluation should treat clean-environment executability as a first-class metric a...

Bhanu Prakash Vangala, Ashish Gehani, Tanu Malik · 0 citations
Open access 2026

CCGMAS: A Multi-Agent Framework for Cross-Platform Go Code Generation via Requirement-Centered Semantic Modeling and Feedback-Driven Verification

CCGMAS enables more explicit semantic alignment across platforms by introducing requirement documents as an intermediate semantic layer and incorporating platform residue modeling, and a feedback-driven refinement loop is designed to iteratively correct errors at different stages, improving both functional correctness...

Xiao Zhang, Bo Yang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.