Knowledge component-constrained diagnostic prompting for automated knowledge gap detection
Abstract
Introductory programming courses face challenges to give scalable feedback on students’ understanding of core programming concepts. Automated grading shows whether code passes its test cases, but not the concept gaps behind the errors. Knowledge Tracing models target those concepts, but they demand extensive historical data, machine-learning expertise, and Graphical Processing Unit (GPU) hardware that many instructors lack. This thesis introduces Knowledge Component-Constrained Diagnostic Prompting (KCDP), a framework that diagnoses programming gaps through prompt design. KCDP directs a generic commercial Large Language Model to the concepts each problem is designed to test, requires it to reason through the code before naming any gap. It then maps each root cause to a set of Knowledge Components (KCs) the instructor selects for the course. Because KCDP relies only on prompting, it can be deployed across commercial LLM platforms using standard access and an existing KC taxonomy. KCDP was evaluated against two human raters. Human raters and the model were all restricted to the same KC vocabulary, so their diagnoses can be measured against each other. Using Google's Gemini 2.5 Flash, KCDP reached an F1 of 0.839 against the human agreement ceiling of 0.885 (94.8% of human agreement) with a Cohen's κ of 0.557 against a human-human κ of 0.669. KCDP results held on when ran on a second, unrelated model (DeepSeek), suggesting it is the prompt design, not the model, that is doing the work. When KCDP flagged a struggling student as weak in a specific concept, that student failed the next problem testing the same concept 77% of the time, against a 26.4% chance rate, while strong students were rarely flagged.