Skip to content

Fairness-aware artificial intelligence tutoring for multilingual learners: evaluating adaptive feedback across L1 and proficiency groups

Sep 2026 · Frontiers in Artificial Intelligence · 0 citations · 109 references
Intelligent Tutoring Systems and Adaptive Learning

Abstract

Recent advances in large language models (LLMs), such as GPT-4, have created new opportunities for intelligent computer-assisted language learning, particularly through personalized feedback for multilingual English learners. Building on research in automated writing evaluation (AWE), this study developed an equity-aware AI tutoring framework that integrates GPT-generated feedback with rule-based simulated feedback for learner spelling errors. The framework draws on annotated TOEFL-Spell data enriched with learner first-language (L1) and English proficiency information to examine feedback quality, accessibility, personalization, and potential subgroup disparities. A verified matched analysis of 32 complete GPT-simulated feedback pairs was conducted using feedback length, Flesch Reading Ease, BERTScore, non-parametric subgroup tests, paired comparisons, effect-size estimates, and exploratory regression models controlling for selected input characteristics. GPT-generated feedback demonstrated substantial semantic correspondence with simulated feedback, achieving an overall BERTScore precision of 0.8370, recall of 0.8742, and F1 score of 0.8552. GPT feedback was significantly longer than simulated feedback ( M = 608.12 vs. 218.69 characters, p < 0.001) and, contrary to the preliminary analysis, was also significantly more readable on average ( M = 75.62 vs. 64.64, p < 0.001). GPT feedback length did not differ significantly across L1 ( p = 0.143) or proficiency groups ( p = 0.422), while GPT readability showed a significant omnibus difference across proficiency levels ( p = 0.047); however, no pairwise comparison remained significant after Holm correction. BERTScore F1 was stable across both L1 ( p = 0.775) and proficiency groups ( p = 0.500). Exploratory controlled analyses likewise provided no consistent evidence that L1 independently predicted feedback readability or semantic alignment after accounting for prompt and error characteristics. Accordingly, observed subgroup variation is interpreted as descriptive heterogeneity rather than definitive evidence of algorithmic or sociolinguistic bias. Because L1 and proficiency were not fully crossed in the analytical sample, their independent effects could not be completely disentangled. A Gradio-based interface ( https://huggingface.co/spaces/Oluyori/equity-ai-tutor ) was additionally developed to demonstrate the practical deployment of the tutoring framework. The findings demonstrate LLM-based feedback potential for semantically aligned and accessible language support while emphasizing the need for larger, balanced samples, learner-centered evaluation, and multidimensional fairness assessment before claims regarding equitable educational performance can be made.

Read PDF

Similar papers

#computer vision Open access Jun 2016

Software Development in Startup Companies: The Greenfield Startup Model

The results are packaged in the Greenfield Startup Model (GSM), which explains the priority of startups to release the product as quickly as possible, and the need to shorten time-to-market, by speeding up the development through low-precision engineering activities.

Carmine Giardino, Nicolò Paternoster, M. Unterkalmsteiner et al. · 178 citations · ⚡14
#computer vision Open access Oct 2016

Software Startups - A Research Agenda

Software startup companies develop innovative, software-intensive products within limited timeframes and with few resources, searching for sustainable and scalable business models.

M. Unterkalmsteiner, P. Abrahamsson, Xiaofeng Wang et al. · 157 citations · ⚡17
#machine learning Review Open access Oct 2016

“Failures” to be celebrated: an analysis of major pivots of software startups

This study conducts a case survey study based on the secondary data of the major pivots happened in 49 software startups, and demonstrates that customer need pivot is the most common among all pivot types.

Sohaib Shahid Bajwa, Xiaofeng Wang, Anh Nguyen-Duc et al. · 127 citations · ⚡15
#computer vision Review Open access May 2015

A survey study on major technical barriers affecting the decision to adopt cloud services

The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.

Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al. · 111 citations · ⚡8
#computer vision Conference Open access Dec 2013

Affordable and Energy-Efficient Cloud Computing Clusters: The Bolzano Raspberry Pi Cloud Cluster Experiment

The ongoing work building a Raspberry Pi cluster consisting of 300 nodes is presented, with potential use cases being an inexpensive and green test bed for cloud computing research and a robust and mobile data center for operating in adverse environments.

P. Abrahamsson, S. Helmer, Nattakarn Phaphoom et al. · 110 citations · ⚡7
#computer vision Book Open access Mar 2017

On the Unhappiness of Software Developers

The results indicate that software developers are a slightly happy population, but the need for limiting the unhappiness of developers remains, and 219 factors representing causes of unhappiness while developing software are identified.

D. Graziotin, Fabian Fagerholm, Xiaofeng Wang et al. · 84 citations · ⚡6

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.