Skip to content

Investigating first-language bias in LLM-based automated essay scoring: A cross-prompt evaluation of an open-weight AI-model on TOEFL essays

Jul 2026 · arXiv.org · Vol abs/2607.14605 · 0 citations · 27 references
Computer Science

TL;DR

This study examines the cross-prompt generalization and first-language (L1) scoring effects of a LoRA-adapted open-weight large language model (Gemma-3-27B-it) applied to automated essay scoring and presents the first large-scale L1 fairness analysis of a fine-tuned open-weight LLM for automated essay scoring.

Abstract

This study examines the cross-prompt generalization and first-language (L1) scoring effects of a LoRA-adapted open-weight large language model (Gemma-3-27B-it) applied to automated essay scoring. Using the identical model and inference configuration reported in"AiAWE: An Open-Source LLM Automated Writing Evaluation System Using LoRA-Adapted Instruction-Tuned Models"(Gayed, 2026), which was fine-tuned on 480 argumentative essays from two prompts, we evaluate scoring accuracy on the full TOEFL11 corpus: 12,100 essays written by test-takers from 11 first-language backgrounds across eight prompts, none of which were seen during training. The model's raw scores (0.5-5.0) are mapped to the same three proficiency bands (low, medium, high) used by ETS, enabling direct comparison. The model achieved an overall band agreement of 77.79% and a quadratic weighted kappa of 0.702, with adjacent-band agreement of 99.98%. Accuracy was stable across all eight unseen prompts, with no advantage for prompts thematically related to the training data, indicating robust cross-prompt generalization. However, the model exhibited a systematic, L1-linked scoring offset. Within every proficiency band, essays from European-language backgrounds received consistently higher scores than essays from East-Asian-language backgrounds, a pattern not attributable to the composition of the fine-tuning data. This is the first large-scale L1 fairness analysis of a fine-tuned open-weight LLM for automated essay scoring.

View source

Similar papers

Review Open access Aug 2026

Large language models as judges: recent advances in LLM-based evaluation, critique, preference modeling, and feedback for text and code

This survey provides a comprehensive overview of recent advances in LLM-based evaluation, covering techniques, applications, and challenges across domains, with future directions emphasizing standardized protocols, uncertainty estimation, and human–AI collaboration.

M. Nadăş · 0 citations
Preprint Aug 2026

ArguLens: An Open-Source System for Automated Essay Scoring and Label-Aware Feedback Generation

Most automated essay scoring (AES) systems output a single holistic score without interpretable evidence and rely on closed APIs that introduce data privacy and cost barriers. We present ArguLens, an opensource, locally deployable system that decomposes AES into three decoupled components: a discourse-move classifier (...

Weiran Wang, Hong-Xiang Shi, Huitao Tang et al. · 0 citations
Open access Jul 2026

Evaluating AI-Based Automated Essay Scoring Through Signal Detection Theory: Beyond Aggregate Agreement Metrics.

The rapid adoption of Large Language Models (LLMs) in educational assessment has reshaped scoring practices, yet evaluation remains tethered to aggregate reliability metrics like Quadratic Weighted Kappa, which obscure discrimination and rater effects. This study applies Signal Detection Theory to evaluate eight state-...

Xiaoliang Zhou · 0 citations
#natural language process... Preprint Sep 2026

Right Tool, Right Job: Native-Language Evaluation, Tokenizer Sensitivity, and Methodological Findings from a French-Only BabyLM

We submit M\'eTRON-FR, a 125M GPT-2 pretrained on 92.47M words of French, to the BabyLM 2026 Strict track. It scores 85.97 +/- 0.17% on QFrBLiMP (a native Quebec-French benchmark of grammatical minimal pairs) and 62.80% on the BabyLM-weighted leaderboard. A cross-lingual GLUE (General Language Understanding Evaluation)...

A. Wasserman, David Beauchemin · 0 citations
Open access Jul 2026

Effect of large language model assistance on undergraduate art history question-answering performance: a randomized crossover pilot study

These pilot findings suggest that supervised LLM assistance may support art history question-answering and explanatory feedback, and future studies should validate these findings in larger cohorts, assess delayed learning retention, and examine open-ended, image-based, and higher-order art history tasks before curricul...

Yunting Zhang, Fan Zhang, Zi-Li Zhang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.