While the computer-generated feedback was broadly considered useful by students, student engagement patterns were markedly different in the solo setting, with students demonstrating reluctance to use the interface's built-in help features and tending to internalize failure in unproductive ways counter to the intention of a formative learning environment.
Abstract
PER has consistently demonstrated the effectiveness of small-group tutorials in helping students develop conceptual understanding and fluency, but instructor uptake is limited by resource constraints. To test the effectiveness of out-of-class tutorials using computer-generated feedback as an instructor-friendly alternative, we conducted think-aloud interviews with students in a quantum computing course who were randomly assigned to either a traditional validated small-group, pencil-and-paper tutorial on tensor products, or a solo computerized adaptation thereof. We found that while the computer-generated feedback was broadly considered useful by students, student engagement patterns were markedly different in the solo setting, with students demonstrating reluctance to use the interface's built-in help features and tending to internalize failure in unproductive ways counter to our intention of a formative learning environment. We discuss implications for curriculum design and directions for future research that may help to answer the longstanding question in PER of why tutorials work so well.
Providing timely and actionable feedback on student writing is a known challenge in large English as a Second Language (ESL) classrooms, where instructor workload often limits the depth and consistency of feedback. This article presents the development and calibration of the COM Essay Assessor, a rubric-based generative artificial intelligence (GenAI) tool designed to support formative feedback while retaining instructor oversight. Grounded in social constructivist theory and formative assessment research, the tool was calibrated using archival student essays and faculty feedback to reflect course-specific evaluation practices. Rather than functioning as an autonomous evaluator, the tool operates within a human-in-the-loop workflow in which instructors review and authorize AI-generated feedback before it reaches students. The article situates this approach within broader work on automated writing evaluation and AI-supported learning and reflects on the opportunities and challenges of integrating GenAI into large writing programs. It argues that rubric-aligned, human-mediated systems can help sustain feedback processes in resource-constrained contexts while preserving pedagogical intent and instructor judgment.
Juhi Bansal· The International Journal of...· 0 citations
The present paper focused on examining the application of GPT‑5.5 as an automated corrective feedback tool for improving the writing skills of 60 Omani pre‑intermediate EFL students enrolled in a foundation program. A quasi‑experimental pretest–posttest design was used, and the participants were divided into two groups, namely an experimental group and a control group, each receiving instruction on advantages‑and‑disadvantages essays. While both groups followed the same syllabus and classroom procedures, only the experimental group utilized GPT‑5.5 AI tool for iterative feedback and revision, supplemented by teacher explanation and clarification. The analytic scoring of the pretest and posttest compositions demonstrated that both groups made progress, yet the experimental group scored substantially better on the posttest, with a large effect size. The results imply that the use of AI‑mediated feedback, when systematically integrated with teacher support, can be an effective and scalable way to advance EFL learners’ writing skills in similar contexts.
Abdullah Al Maqbali, O. Taga· Language, Technology, and So...· 0 citations
This case study examines how middle school students in a rural, Indigenous-serving school engaged with artificial intelligence (AI)-generated writing feedback during a classroom journal-writing activity. Drawing on think-aloud sessions with four focal students, supplemented by researcher observation notes and brief post-activity student and teacher surveys, the study focuses on observable patterns in how students engaged with, responded to, and used automated feedback under real classroom conditions. Analysis identified four interrelated dimensions: (1) AI feedback as actionable guidance for revision, (2) rubric-aligned feedback as a structure for focused revision, (3) usability and access constraints as central conditions shaping engagement, and (4) variation in students’ use of interactive AI features and in their experiences of feedback readability. The findings suggest that AI-generated feedback supported targeted revision when it is accessible, interpretable, and aligned with classroom assessment criteria. At the same time, technical disruptions, feedback length, linguistic complexity, and local infrastructural conditions significantly shaped students’ uptake of the tool. By focusing on classroom implementation rather than performance outcomes, this study contributes empirically grounded insights into the situated affordances and limitations of AI-supported writing feedback in an underrepresented K-12 context.
A. Tzirides, Michele Galla, B. Cope et al.· Ubiquitous Learning An Inter...· 0 citations
Providing high-quality feedback on student work is essential for learning, yet delivering such feedback at scale remains challenging. In this paper, we focus on feedback for open-ended short answer questions in introductory programming, with the goal of nudging students toward success on reattempts without revealing the correct answer. We develop a five-criteria rubric grounded in educational literature for evaluating feedback quality: (1) acknowledging correct portions of the student answer, (2) identifying at least one flaw (if present), (3) providing actionable guidance for improvement, (4) maintaining appropriate concealment of the answer, and (5) using an appropriate conversational tone. Using this rubric, we compare feedback generated by a frontier LLM (OpenAI o1) to feedback from nine teaching assistants across 90 student responses, with three researchers and an LLM independently scoring all feedback. Our results show that while one TA often produced the best feedback, the LLM demonstrated consistently higher average performance than TAs, as evaluated by humans. However, we also uncover significant self-preference bias when using LLMs to evaluate feedback quality: the LLM systematically rated its own outputs higher than human experts did. This bias, which research suggests persists even in cross-model evaluation, raises important methodological concerns for researchers employing LLM-based evaluation. We provide detailed characterization of both TA and LLM performance, analyze sources of variance in TA feedback quality, and discuss implications for deploying LLM-generated feedback in educational settings.
Binglin Chen, Rajarshi Haldar, Max Fowler et al.· 0 citations
This study investigated how college students’ use of ChatGPT has evolved from a tool for simply getting answers to one that supports motivation and learning engagement. Using a novel survey instrument based on Self-Determination Theory, we measured six types of motivation, three intrinsic, two extrinsic, and amotivation, across two cohorts of students (2023 and 2025). Results revealed statistically significant increases in both intrinsic and extrinsic motivation over time. While early use of ChatGPT focused on convenience and task completion, students increasingly reported using it to overcome mental blocks, build momentum, and stay engaged with their academic work. Follow-up analyses of repeat participants and comparisons by class standing suggest these trends reflect broader shifts in student engagement, not just cohort or experience differences. Cluster analysis revealed distinct motivational profiles, highlighting variation in how students incorporate ChatGPT into their academic routines. These findings suggest that as generative artificial intelligence becomes more familiar, accessible, and sophisticated, students are integrating it more deeply into their learning processes, not just for answers, but for sustained motivation and support.
Jayne Ann Harder, C. E. Klehm, A. Lang et al.· Artificial Intelligence and...· 0 citations
Open-ended short responses can reveal more about student understanding than selectedresponse items yet scoring them during live instruction is rarely feasible. This paper presents a bring-your-own-device (BYOD) mobile classroom system that combines rubric-guided large language model (LLM) scoring of short open-ended answers, real-time teacher analytics through a classroom dashboard, and evidence-based post-quiz feedback generation. The system includes a teacher-reviewed workflow in which instructors can inspect model outputs and, when needed, adjust scores or feedback during formative classroom use. In the reported evaluation, however, teacher override was not applied or analyzed; the agreement metrics compare the system’s initial rubric-guided outputs with expert reference scores. The system was evaluated in three classroom groups (N = 60 students; 480 question-level scoring cases). Three expert raters participated, each assigned to one classroom group, so each response received one expert reference score. In this test, system scores showed strong but not complete agreement with expert reference scores (QWK = 0.887; MAE = 0.089; RMSE = 0.121). Mastery-label accuracy reached 78.3%. Expert raters gave positive scores to the generated feedback reports for groundedness (M = 4.78), specificity (M = 4.37), and actionability (M = 4.57) on a 1–5 scale. Within the limits of this small-scale classroom evaluation, the findings suggest that rubric-guided LLM scoring may support formative assessment in mobile classrooms, provided teacher oversight is maintained for borderline cases.
Abylay Yerniyazov, Bakyt Bakayeva, Zhalgasbek Iztayev et al.· International Journal of Int...· 1 citation