A five-level scale is developed, validated against human annotations, that characterizes responses according to the degree of direct assistance LLMs provide when students use them as tutors in authentic learning interactions and provides an empirical baseline for LLM assistance in tutoring interactions and a measurement framework for evaluating how alternative tutoring designs change that assistance.
Abstract
Students increasingly use LLMs as tutors for coursework and problem solving. Little is known about the level of assistance LLMs provide when students use them as tutors in authentic learning interactions. This matters because tutoring responses can differ substantially in how directly they help students complete a task. We operationalize this dimension as scaffolding level and develop a five-level scale, validated against human annotations, that characterizes responses according to the degree of direct assistance they provide. We apply the scale to 14,637 LLM responses from 203 students in a university AI course. Responses are overwhelmingly concentrated at high levels of assistance, with more than 95% classified as either Explaining or Solving. Scaffolding level is systematically associated with students'subsequent conversational behavior, but provides little additional predictive information about performance on three subsequent exams beyond prior achievement and dialogue behavior. These findings provide an empirical baseline for LLM assistance in tutoring interactions and a measurement framework for evaluating how alternative tutoring designs change that assistance.
The results suggest that learner- and curriculum-aware alignment may matter more for effective tutoring than model category alone, and that such alignment is both measurable and improvable.
Benjamin Barlog, Hudson Craig, Ze-Dong Peng· IEEE International Conferenc...· 0 citations
Simulated students generated by computational models provide a practical way to evaluate tutoring strategies and pedagogical approaches used by human and AI tutors. However, such simulations should account for disengaged behaviors, including gaming the system, wheel-spinning, and off-task behavior, because tutors may n...
A simulation-based textual analysis of prompt design evaluates a frontier large language model as a tutor across 60 scripted sessions on a single topic and point to dynamic, dialogue-aware prompting alongside explicit SRL scaffolding.
Kendall Hartley, Fabiola Sáez-Delgado, Javier Mella-Norambuena· Future Internet· 0 citations
EduClaw-Bench is introduced, a benchmark that places an agent tutor in a continuous 30-day relationship with a simulated learner grounded in knowledge tracing (KT), whose knowledge-concept mastery, from a KT model trained on real-student data, drives its answers and is probed for learning gain across 55 scenarios.
Unggi Lee, Sookbun Lee, Yeil Jeong et al.· 0 citations
The results suggest that the pedagogical behavior of AI tutors may not be easily steered through system prompts alone: embedding established SRL and CE frameworks did not produce detectable improvements on any preregistered outcome in a large, ecologically valid deployment.
Maximilian Georg Barth, Sverrir Thorgeirsson, K. Etemadi et al.· International Computing Educ...· 0 citations
This work presents StudentSim, a training framework that turns sparse per-student data into individualized simulators through pooled training followed by per-student specialization, and introduces StudentSimEval, a standardized protocol covering 60 students across chess, second-language English writing, and mathematics...
Ke Yang, Chenglong Wang, Michel Galley et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.