Aug 2026· International Computing Education Research Workshop· 0 citations· 49 references
Computer Science
TL;DR
The results suggest that the pedagogical behavior of AI tutors may not be easily steered through system prompts alone: embedding established SRL and CE frameworks did not produce detectable improvements on any preregistered outcome in a large, ecologically valid deployment.
Abstract
Background. Large language models are increasingly deployed as tutors in introductory programming courses, yet evidence that they actually improve learning remains thin, and their tendency to shortcut productive struggle raises concerns about pedagogical harm. Self-regulated learning (SRL) and cognitive engagement (CE) frameworks offer a principled way to address this, but whether embedding them in system prompts actually changes how students learn is an open question. Objectives. We investigated whether AI tutors guided by SRL and CE frameworks affect conceptual understanding, perceived usability, cognitive load, student engagement, and other outcomes when compared to a pedagogically constrained baseline tutor in CS1. Methods. We conducted a preregistered, three-armed crossover study over six weeks of authentic coursework, comparing a baseline AI tutor against two SRL and CE-guided tutors designed to model Zimmerman’s cyclical model, which scaffolds planning, monitoring, and reflection, and Chi’s ICAP framework, which promotes progressively deeper forms of cognitive engagement. We assessed outcomes using post-exercise surveys, conceptual multiple-choice questions, in-platform ratings, and coded responses from the interaction logs, and analyzed the quantitative measures with mixed-effects models. Findings. On the four preregistered confirmatory measures, we found no statistically significant differences between conditions. Non-confirmatory analysis showed that students spent significantly more time on task, wrote longer messages, and produced more constructive contributions when interacting with SRL and CE tutors. Furthermore, the relationship between cognitive load and quiz performance differed significantly by agent type. Implications. Our results suggest that the pedagogical behavior of AI tutors may not be easily steered through system prompts alone: embedding established SRL and CE frameworks did not produce detectable improvements on any preregistered outcome in a large, ecologically valid deployment. Rather than prescribing a single tutoring strategy, future designs may benefit from giving students greater agency over the kind of help they receive, allowing them to choose between scaffolded and more direct support based on their own needs.
A simulation-based textual analysis of prompt design evaluates a frontier large language model as a tutor across 60 scripted sessions on a single topic and point to dynamic, dialogue-aware prompting alongside explicit SRL scaffolding.
Kendall Hartley, Fabiola Sáez-Delgado, Javier Mella-Norambuena· Future Internet· 0 citations
Large language models (LLMs) are increasingly used as on-demand conversational learning assistants, but they typically do not adapt explanations to a student’s background unless explicitly prompted. We present the Personalized Learning Assistant Interface (PLAI), a web-based prototype that generates explanations from l...
Furkan Ali Yurdakul, Yi-Man Wu, Maria Torres Vega et al.· Message Understanding Confer...· 0 citations
This study examines the role of perceived desirable difficulty in AI tutoring systems and explains how productive learning challenge influences learner engagement, productive struggle orientation, intrinsic motivation, knowledge retention confidence and continued use intention. Although AI tutoring is increasingly used...
Sumit Khanal, Abhijeet Sen Gupta· Australian Journal of Artifi...· 0 citations
The results suggest that learner- and curriculum-aware alignment may matter more for effective tutoring than model category alone, and that such alignment is both measurable and improvable.
Benjamin Barlog, Hudson Craig, Ze-Dong Peng· IEEE International Conferenc...· 0 citations
A five-level scale is developed, validated against human annotations, that characterizes responses according to the degree of direct assistance LLMs provide when students use them as tutors in authentic learning interactions and provides an empirical baseline for LLM assistance in tutoring interactions and a measuremen...
Suhyeon Lee, Juneha Baek, Jaehyeong Park et al.· 0 citations
Students often struggle to effectively self‐regulate their learning strategies during writing tasks and typically require external support from teachers, peers or AI tools to overcome these challenges. While emerging research explores how generative AI can dialogically interact with students to support the writin...
Sehrish Iqbal, Jelena Jovanović, Yizhou Fan et al.· Journal of Computer Assisted...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.