Similar papers
A Systematic Synthesis of Research on Popular Corpus Tools in English Language Learning and Teaching
It is well-established that second/foreign language learning requires exposure to meaningful input. Thanks to advancing technologies, one way to achieve this in language teaching contexts is via online corpus tools. Given the potential pedagogical gains these tools offer for language skills (e.g., vocabulary, grammar, pronunciation, and writing), this study aims to synthesise the findings of research on five popular corpus tools – Sketch Engine, SkELL, PlayPhraseMe, Fraze.it, and CorpusMate. Following the PRISMA guidelines for systematic reviews, this study included five research papers from the Web of Science and Scopus databases until 2024. Applying qualitative and quantitative content analyses, the study found that corpus consultation promotes language skills and learner autonomy, supports contextually appropriate language use, and improves sensitivity to authentic language structures. Despite the promising outcomes, the scant number of studies analysed with some methodological issues and contextual constraints undermines the generalisability of the results. The study underscores the urgent need for more longitudinal, comparative, and large-scale research on using these corpus tools to harness them fully in language education. It also proposes that when systematically implemented, corpus tools may serve as powerful resources for language learning and teaching.
The KSAUHS learner corpus
The KSAU-HS Learner Corpus is a longitudinal corpus of EFL tertiary writing that complies with the FAIR principles. Collection began in 2022 and captures writing development during a period of emerging language technologies (2022–24). The corpus contains over 856,907 tokens across 2,387 texts produced by 157 preparatory year university students, within a CEFR-aligned program with an instructional range of approximately A2-B2. Texts span four trimesters and include rhetorical modes such as cause-and-effect, argumentation, and summarisation. Metadata includes years of English schooling, other languages spoken, and preferred reference tools. The corpus enables research into writing development to inform EAP pedagogy and assessment. Avenues for investigation include lexico-grammatical development, cross-linguistic influence, individual differences, and the impact of task conditions and language technologies. This resource promises data-driven insights into the textual features and factors that characterise EFL writing proficiency, and plans are in place to expand its size, representation, and accessibility.
Data-driven learning in vocabulary instruction for South Korean learners of English
Corpus-based tools and techniques not only facilitate the description of learners' language but can also be used in educational settings to design targeted classroom activities. This approach is known as data-driven learning (DDL). Access to corpus data enables learners to observe and analyze patterns in real language use, addressing their specific needs and fostering learning autonomy. While several studies have examined the effectiveness of DDL in teaching various language skills, few have investigated its impact highlighting the implications for specific cultural groups. Adopting a cumulative knowledge building perspective, this study systematically synthesizes and builds upon previous empirical research to advance our understanding of the effectiveness of DDL for South Korean learners of English. The purpose is paper to (1) survey different DDL activities piloted in vocabulary instruction across various English language teaching contexts in Korea; (2) determine the effectiveness of DDL for vocabulary instruction for this demographic; and (3) explore Korean learners' attitudes towards DDL. To do so, empirical studies were systematically identified using the Korea Citation Index with the keywords “data-driven learning,” “corpus-based,” “vocabulary,” “lexis,” “Korea,” and “Korean.” The results indicate that DDL is generally welcomed by and effective for students in this demographic. By building on cumulative findings, this study provides empirical foundation for curriculum development, classroom practices, and teacher training. Tailoring DDL activities to the Korean context can maximize their effectiveness addressing learners' unique linguistic challenges.
Learning Sequences and Events in Individual, Pair and Group Data-Driven Learning
Abstract Previous research has shown that the learning process in data-driven learning (DDL) is cyclical, iterative, and non-linear, even when tasks are designed around seemingly linear pedagogical frameworks such as Illustration–Interaction–Intervention–Induction (Carter, Ronald, and Michael McCarthy. 1995. “Grammar and the Spoken Language.” Applied Linguistics 16 (2): 141–58; Flowerdew, Lynne. 2009. “Applying Corpus Linguistics to Pedagogy: a Critical Evaluation.” International Journal of Corpus Linguistics 14 (3): 393–417). Building on this work, the present study investigates how different DDL configurations shape both learning outcomes and processes in academic reading comprehension. Drawing on constructivism, sociocultural theory, and the noticing hypothesis, the study adopts an established analytical framework distinguishing between seven macro-level learning sequences and twelve micro-level learning events in DDL. French-speaking Master’s students enrolled in a linguistics programme participated in a six-week hands-on DDL intervention targeting move analysis of English research article introductions (Swales, John M. [1981] {2011}. Aspects of Article Introductions. Michigan: Michigan Classics Edition). Learners worked individually, in pairs, or in groups while consulting the British Academic Written English corpus via Sketch Engine. Learning outcomes were evaluated using pre- and post-tests analysed with Bayesian linear mixed-effects regression, while learning processes were examined through video-based observation and annotation in ELAN. Results indicate that DDL led to overall gains in academic reading comprehension, with the magnitude of these gains varying by configuration: Pair DDL was most likely to support improvement, whereas Group showed weaker and more variable outcomes. Behavioural analyses reveal Pair offered a balance between scaffolding and efficiency, whereas Individual relied heavily on self-regulation and Group incurred substantial coordination cost. These findings suggest that balancing learner autonomy and peer collaboration may be a relevant consideration in DDL design.
Artificial Intelligence in English Language Teaching and Learning: A Scoping Review of Intelligent Computer-Assisted Language Learning (2015–2025)
This scoping review primarily aims to synthesize empirical research on artificial intelligence in English language teaching and learning published from 2015 to 2025. The main question investigates how the integration, applications, and pedagogical roles of AI have evolved over the past decade. The significance of this study is that it uses Intelligent Computer-Assisted Language Learning as an interpretive lens to make sense of a rapidly shifting field, offering a framework to help educators navigate modern generative tools. Following Preferred Reporting Items for Systematic Reviews and Meta-Analyses Extension for Scoping Reviews (PRISMA-ScR) guidance, 129 empirical studies were identified and analyzed using descriptive mapping and thematic analysis. The main findings indicate three overlapping evolutionary phases: an early system construction phase focused on tutoring, a mobile-and-voice phase emphasizing speech practice, and a generative phase dominated by large language models. The evidence base remains heavily concentrated in higher education, where AI frequently acts as a tutor, practice partner, or co-writer. For further use, this study recommends adopting teacher-mediated task designs, shifting assessments to focus on the learning process, and prioritizing longitudinal research in primary and under-resourced educational settings.
Direct Use of Corpora in Japanese Language Education: Implementing DDL at the Upper-Intermediate Level
Large-scale corpora offer access to authentic language and represent a valuable resource for Japanese language education. In current practice, they are mostly used indirectly by instructors, while direct learner engagement with corpora (data-driven learning, DDL) remains relatively limited in Japanese courses, partly due to time constraints and the complexity of corpus tools. This article reports on a four-year project conducted at Ca’ Foscari University of Venice. The course integrated direct corpus consultation through NINJAL corpora accessed via Chūnagon, with the dual aim of supporting language development and fostering research-oriented skills for working with Japanese primary sources. The study provides a qualitative evaluation based on students’ corpus-based presentation projects and anonymous student comments, highlighting both perceived benefits and critical issues, thus offering practical insights into the feasibility of DDL in Japanese language education.