Sep 2026· Frontiers in Artificial Intelligence· 0 citations· 109 references
Intelligent Tutoring Systems and Adaptive Learning
Abstract
Recent advances in large language models (LLMs), such as GPT-4, have created new opportunities for intelligent computer-assisted language learning, particularly through personalized feedback for multilingual English learners. Building on research in automated writing evaluation (AWE), this study developed an equity-aware AI tutoring framework that integrates GPT-generated feedback with rule-based simulated feedback for learner spelling errors. The framework draws on annotated TOEFL-Spell data enriched with learner first-language (L1) and English proficiency information to examine feedback quality, accessibility, personalization, and potential subgroup disparities.
A verified matched analysis of 32 complete GPT-simulated feedback pairs was conducted using feedback length, Flesch Reading Ease, BERTScore, non-parametric subgroup tests, paired comparisons, effect-size estimates, and exploratory regression models controlling for selected input characteristics.
GPT-generated feedback demonstrated substantial semantic correspondence with simulated feedback, achieving an overall BERTScore precision of 0.8370, recall of 0.8742, and F1 score of 0.8552. GPT feedback was significantly longer than simulated feedback (
M
= 608.12 vs. 218.69 characters,
p
< 0.001) and, contrary to the preliminary analysis, was also significantly more readable on average (
M
= 75.62 vs. 64.64,
p
< 0.001). GPT feedback length did not differ significantly across L1 (
p
= 0.143) or proficiency groups (
p
= 0.422), while GPT readability showed a significant omnibus difference across proficiency levels (
p
= 0.047); however, no pairwise comparison remained significant after Holm correction. BERTScore F1 was stable across both L1 (
p
= 0.775) and proficiency groups (
p
= 0.500). Exploratory controlled analyses likewise provided no consistent evidence that L1 independently predicted feedback readability or semantic alignment after accounting for prompt and error characteristics.
Accordingly, observed subgroup variation is interpreted as descriptive heterogeneity rather than definitive evidence of algorithmic or sociolinguistic bias. Because L1 and proficiency were not fully crossed in the analytical sample, their independent effects could not be completely disentangled. A Gradio-based interface (
https://huggingface.co/spaces/Oluyori/equity-ai-tutor
) was additionally developed to demonstrate the practical deployment of the tutoring framework. The findings demonstrate LLM-based feedback potential for semantically aligned and accessible language support while emphasizing the need for larger, balanced samples, learner-centered evaluation, and multidimensional fairness assessment before claims regarding equitable educational performance can be made.
The results are packaged in the Greenfield Startup Model (GSM), which explains the priority of startups to release the product as quickly as possible, and the need to shorten time-to-market, by speeding up the development through low-precision engineering activities.
Carmine Giardino, Nicolò Paternoster, M. Unterkalmsteiner et al.· IEEE Transactions on Softwar...· 178 citations· ⚡14
Software startup companies develop innovative, software-intensive products within limited timeframes and with few resources, searching for sustainable and scalable business models.
M. Unterkalmsteiner, P. Abrahamsson, Xiaofeng Wang et al.· e-Informatica Software Engin...· 157 citations· ⚡17
This study conducts a case survey study based on the secondary data of the major pivots happened in 49 software startups, and demonstrates that customer need pivot is the most common among all pivot types.
Sohaib Shahid Bajwa, Xiaofeng Wang, Anh Nguyen-Duc et al.· Empirical Software Engineeri...· 127 citations· ⚡15
The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.
Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al.· Journal of Systems and Softw...· 111 citations· ⚡8
The ongoing work building a Raspberry Pi cluster consisting of 300 nodes is presented, with potential use cases being an inexpensive and green test bed for cloud computing research and a robust and mobile data center for operating in adverse environments.
P. Abrahamsson, S. Helmer, Nattakarn Phaphoom et al.· IEEE International Conferenc...· 110 citations· ⚡7
The results indicate that software developers are a slightly happy population, but the need for limiting the unhappiness of developers remains, and 219 factors representing causes of unhappiness while developing software are identified.
D. Graziotin, Fabian Fagerholm, Xiaofeng Wang et al.· International Conference on...· 84 citations· ⚡6
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduOct 8, 2026