Skip to content

Prompt Coach: An Empirical Evaluation of an Agentic Tutor for Learning Prompt Engineering in Software Development

Jul 2026 · arXiv.org · Vol abs/2607.06074 · 0 citations · 35 references
Computer Science

Abstract

Prompt engineering has emerged as a critical yet undertaught skill for software developers, one that traditional learning approaches are ill-equipped to support given its evolving, interactive, and context-dependent nature. In this paper, we introduce Prompt Coach (PC), an agentic tutor that helps developers learn how to craft high-quality code-generation prompts through Socratic guidance embedded in-flow within their IDE. PC evaluates prompt quality across multiple dimensions and surfaces targeted questions to guide self-correction, grounded in the developer's codebase and the behavior of the target LLM. We present an early empirical study with 15 professional developers combining quantitative prompt quality scoring with qualitative perception measures. Participants showed statistically significant improvements after a single 60-minute session, with the largest gains across dimensions commonly overlooked by developers. They also reported strong trust, high adoption readiness, and unanimous agreement that PC improved their prompt-writing skills.

View source

Similar papers

Open access Jul 2026

Prompting for Independent Learning: An Evaluation of Tutoring Behaviors in GenAI

A simulation-based textual analysis of prompt design evaluates a frontier large language model as a tutor across 60 scripted sessions on a single topic and point to dynamic, dialogue-aware prompting alongside explicit SRL scaffolding.

Kendall Hartley, Fabiola Sáez-Delgado, Javier Mella-Norambuena · 0 citations
Book Open access Aug 2026

Steering AI Tutors Through System Prompts: A Crossover Study on Self-Regulated Learning and Cognitive Engagement Scaffolds in CS1

The results suggest that the pedagogical behavior of AI tutors may not be easily steered through system prompts alone: embedding established SRL and CE frameworks did not produce detectable improvements on any preregistered outcome in a large, ecologically valid deployment.

Maximilian Georg Barth, Sverrir Thorgeirsson, K. Etemadi et al. · 0 citations
Review Open access Aug 2026

Human-in-the-Loop LLM Assessment for Programming Education: Design and Empirical Validation

The system attained a System Usability Scale (SUS) score of 88.5 and cut grading time by 87.5 percent, facilitating focused instructor review in LLM-supported programming assessments, facilitating focused instructor review in LLM-supported programming assessments.

A. Ibrahim, Runal Rezkiawan · 0 citations
Review Open access Aug 2026

The COM Essay Assessor

The development and calibration of the COM Essay Assessor is presented, a rubric-based generative artificial intelligence (GenAI) tool designed to support formative feedback while retaining instructor oversight and reflects on the opportunities and challenges of integrating GenAI into large writing programs.

Juhi Bansal · 0 citations
Review Jul 2026

Structured AI Demonstrations and Student LLM Use in Engineering Mechanics: Study Design and Preliminary Results

This manuscript presents a descriptive study design and preliminary findings from an undergraduate engineering mechanics course conducted in Spring 2026, and details a reproducible survey instrument used to capture student AI usage patterns, attitudes, and verification practices, which are subsequently linked to academ...

S. Geng, Helen Lallos-Harrell, Jiya Ashar et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.