Evaluating an AI-scaffolded intervention for L2 vocabulary learning: affordances, constraints, and pedagogical implications
Abstract
This exploratory mixed-methods study describes vocabulary-score changes in two pre-existing Chinese language classes: one received a ChatGPT-scaffolded instructional approach and the other received traditional vocabulary instruction. Sixteen international undergraduates (eight per class) participated in the six-week pretest-posttest study. At the student level, the ChatGPT class showed a larger observed mean change on a composite semantic-syntactic measure, whereas observed mean phonetic changes were similar across classes. Because instructional condition was completely confounded with class membership, these class-level contrasts cannot establish a treatment effect. Interviews with six self-selected volunteers indicated perceived benefits and limitations of ChatGPT-scaffolded instruction. The findings are preliminary, class-specific observations that may inform stronger future research.