Explanation First Prompting for Reviewing AI Programming Answers in Undergraduate Courses
Abstract
Undergraduate students can now obtain complete programming answers from generative artificial intelligence systems in seconds, but many of those answers are accepted with little checking. This paper examines whether explanation-first prompting makes such answers easier to review rather than whether it improves student learning directly. Ten short Python tasks typical of undergraduate programming courses were posed to a GPT-5-based coding assistant under two conditions: a directanswer prompt and an explanation-first p rompt t hat required task restatement, edge-case identification, a pproach explanation, and a brief self-check before code delivery. The resulting 20 answers were scored by the author with a four-dimension rubric covering task alignment, correctness and edge cases, reasoning transparency, and efficiency and c ode q uality. T he explanationfirst c ondition a chieved a h igher m ean s core t han t he directanswer condition (6.3 vs. 5.0 on an 8-point rubric) and performed better on nine of ten tasks. The clearest improvement appeared in reasoning transparency (1.9 vs. 1.1), followed by task alignment (1.9 vs. 1.6). Gains in correctness and efficiency w ere modest. In this small single-model, single-rater benchmark, asking for explanations before code did not remove defects, establish statistical significance, or demonstrate learning gains. It did, however, expose weak assumptions earlier and reduced requirement drift in the generated artifacts. For programming courses, the result supports a narrow practical claim: if students use AI-generated answers, asking for explanation before code and then reviewing the result with a short rubric is a stronger inspection protocol than asking for code alone.