Skip to content

Vibe Coding: An Experiment with Test-Driven Development

Jul 2026 · arXiv.org · Vol abs/2607.22406 · 1 citation · 78 references
Computer Science

TL;DR

This exploratory study aims to investigate how humans and CLLMs can collaborate as peers through vibe coding, an approach that integrates principles from prompt engineering, agile design, and human-AI co-creation to enhance collaboration.

Abstract

Context: Conversational Large Language Models (CLLMs) can automatically generate code by collaborating with users through natural language. However, poor collaboration can lead to poor quality output. Objective: This exploratory study aims to investigate how humans and CLLMs can collaborate as peers through vibe coding, an approach that integrates principles from prompt engineering, agile design, and human-AI co-creation to enhance collaboration. Method: We designed four interaction models representing different collaboration patterns in the software development process: the solo model (human-only development), the collaborative model (human-CLLM collaboration), the fully automated model (development autonomously performed by a CLLM), and the agentic model (development autonomously performed by the MetaGPT~X platform). Based on these models, we implemented corresponding Test-Driven Development (TDD) workflows using structured prompts and Python scripts. We then conducted a controlled pre-experimental study with TDD professionals to compare the solo and collaborative workflows. In addition, we performed repeated exploratory executions of fully automated and agentic workflows on the same development tasks to obtain complementary evidence. Results: Our findings suggest that the choice of interaction model should depend on the development objective. Agentic workflows are best suited for rapid development and functionally correct production code but may introduce additional implementation complexity. However, they may also introduce additional implementation decisions that are not explicitly required by the functional specifications, resulting in untested decision points. In contrast, collaborative workflows produce higher-quality, better-organized test suites. Conclusions: Our work explored how...

View source

Similar papers

Book Open access Aug 2026

From Design to Code: Exploring LLM-Supported Workflows in Figma Dev Mode

It is suggested that AI-supported workflows can improve developer experience by reducing ambiguity in design interpretation while maintaining the need for human validation and documentation support.

Surbhi Rajpal, Andreas Riener · 0 citations

Look Before You Prompt, and After: Scaffolding Human-AI Collaboration in Software Tutorial Creation

This work designed a tool called dBlocks with the following features: blocks to scope content, a context manager to edit context, and inline execution to verify code, and shows how human-centered design can guide the development of LLM-integrated tools.

Avinash Bhat, Vy Bui, Jin L. C. Guo · 0 citations
Open access Aug 2026

A Multi-Study Evaluation into Generative Artificial Intelligence for Test-Driven Development

GAI4-TDD makes failing tests pass with a success rate of about 90% on the first attempt; improves external/internal quality of software and students’ productivity; and professionals are generally-positive, although some barriers to GAI4-TDD adoption emerged.

P. Cassieri, Simone Romano, Valentina Lenarduzzi et al. · 0 citations
Review Open access Aug 2026

Vibe Coding for Statistical Analysis Using Large Language Models.

While vibe coding has the potential to reduce barriers to data analysis for researchers, this case study demonstrated that it can produce both valid and invalid outputs and that foundational statistical training, knowledge, understanding, and methodological expertise remain paramount when using it.

D. Tolentino, E. Kohout, Paul Boy et al. · 0 citations
Book Open access Oct 2026

Large Language Models Assistance in core Model-Driven Engineering activities

Recent research explores the use of Large Language Models (LLMs) and conversational agents to enable automation and assist modeling tasks across the Model-Driven Engineering (MDE) lifecycle, including design, model management, and evolution. This paper presents a Systematic Literature Review of 96 peer-reviewed studies...

Arianna Fedeli, Maria Teresa Rossi, Ludovico Iovino · 1 citation
#software testing Review Sep 2026

How Developers Discuss Generative AI: A Longitudinal Study of the Visual Studio Code Community

Generative AI tools such as GitHub Copilot, ChatGPT, and coding agents have rapidly become part of everyday software development, yet little is known about how mainstream open source communities discuss them in practice. This paper presents a longitudinal analysis of generative-AI-related discussions in the Visual Stud...

Panida Rumriankit, Akito Monden, Hiroki Inayoshi et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.