Skip to content
Open access

EduFairBench: reproducible evaluation of large language models for educational assessment

Aug 2026 · Frontiers in Artificial Intelligence · Vol 9 · 0 citations · 21 references
Medicine

TL;DR

E EduFairBench provides a reproducible methodology for jointly analyzing predictive performance, robustness, uncertainty, and feedback quality, providing a comprehensive methodological framework for the rigorous evaluation of LLM-based educational assessment systems.

Abstract

Large language models (LLMs) are increasingly used to evaluate open-ended educational responses. However, their performance is often assessed using aggregate metrics that provide limited insight into prediction stability, uncertainty, error patterns, and feedback quality. This study presents EduFairBench, a reproducible evaluation protocol designed to characterize LLM behavior across short-answer assessment and automated essay scoring using open educational benchmarks. The protocol combines repeated inference, majority-vote consolidation, uncertainty estimation, error analysis, and structural evaluation of generated feedback within a unified experimental framework. Experiments were conducted on SciEntsBank, Beetle, and ASAP2, comprising 2,000 student responses and 10,000 independent LLM inferences. The results showed moderate predictive agreement with human assessment while revealing substantial differences between nominal and ordinal evaluation tasks. Repeated inference demonstrated high internal stability across benchmarks, although systematic errors remained in semantically adjacent categories, indicating that prediction consistency does not necessarily imply correctness. Feedback quality varied by task type, with longer textual contexts yielding more specific and pedagogically structured explanations. These findings demonstrate that evaluating educational LLMs requires complementary analyses beyond conventional performance metrics. EduFairBench provides a reproducible methodology for jointly analyzing predictive performance, robustness, uncertainty, and feedback quality, providing a comprehensive methodological framework for the rigorous evaluation of LLM-based educational assessment systems.

Read PDF

Similar papers

Open access Aug 2026

Comparative Evaluation of Large Language Models in Computer Programming Education

A comparative analysis of six LLMs for generating formative feedback on introductory Java programs containing predefined defects under controlled conditions reveals substantial cross-model variation, particularly in multi-defect scenarios.

Melina Najimi, Saba Yazdani, Marzieh Ahmadzadeh · 0 citations
#small language model Open access Sep 2026

Human versus machine

For teachers to effectively use large-language-model(LLM)-based ratings in the formative or summative assessment of texts, it is essential to ensure that such ratings can assess student writing in a valid and reliable manner. This study investigates whether a validated human text-rating procedure (benchmark rating) can be replicated by an LLM-based rating procedure. We tested the replication with two genres of elementary school students’ text—narrative and instructive—using nine LLMs from three providers (OpenAI, Anthropic, Mistral). Each LLM generated three independent scores per text via structured, benchmark-aligned prompts that were then aggregated into a consensus score. Results showed that intrarater reliability was high to excellent, ICC(3, k) ≈ .68–.97, and alignment with human ratings ranged from moderate to strong, ICC(3, 1) ≈ .47–.85, with larger models consistently outperforming smaller ones. Systematic bias patterns emerged, varying by model and genre, indicating a need for calibration. Increasing output token windows and reducing temperature parameters mitigated truncation and schema-related failures. Although LLM-based benchmark ratings can approximate expert judgments and reduce the need for labor-intensive human triple coding, limitations remain regarding cost (for larger models), genre- and task specificity, and sensitivity to text presentation and student grade level—factors that constrain immediate classroom use, particularly for formative feedback.

Afra Sturm, Valentin Unger, Fabian Grünig · 0 citations
Open access Aug 2026

A benchmark dataset with human validation for AI-assisted technical answer evaluation

This study presents DSA-RubricEval, a preliminarily reliability-assessed, rubric-based benchmark dataset intended to support exploratory research on pedagogically meaningful AI-assisted assessment of open-ended technical responses.

J. Sheikh, Hemant Kumar Soni · 0 citations
Open access Aug 2026

RoCulturaMCQ: Building a Benchmark While Learning Statistics

A pilot project in which students in a statistics course within a data science engineering program created culturally diverse multiple-choice questions, generated answers using LLMs, and applied statistical methods to assess model accuracy is presented, supporting a cultural injection hypothesis.

Denis Iorga, Razvan Muntean, Mihai Masala et al. · 0 citations
Open access Aug 2026

GradeDrift-LLM: Measuring Student-History-Induced Score Drift in LLM-Based Automated Grading

Student-history metadata can influence LLM-generated grading scores despite explicit instructions to ignore it, and future LLM-based grading systems should separate answer-based scoring from learner-context-based personalization and validate score invariance under controlled learner-context variations.

Cătălin Anghel, A. Anghel, M. Craciun et al. · 0 citations
Open access Jul 2026

Effect of large language model assistance on undergraduate art history question-answering performance: a randomized crossover pilot study

Introduction Large language models (LLMs) are increasingly used in higher education, yet empirical evidence for their effectiveness in art education remains scarce. This study aimed to evaluate whether LLM assistance could improve undergraduate art history question-answering performance and explanatory support. Methods This study developed the Art History Theory Question Set (AHTQS), comprising 104 single-choice items with Bloom-level annotations, and benchmarked three LLMs (ChatGPT-4o, DeepSeek-V3, and Qwen2.5-Plus). DeepSeek-V3 showed the highest accuracy (96.2%) and lowest observed run-to-run variability and was selected for a randomized crossover pilot study with six undergraduates. The primary outcome was the change in examination accuracy from independent to LLM-assisted answering. A Likert-scale evaluation involving nine students and three instructors was also conducted to assess the clarity and coherence of LLM-generated explanations. Results A one-sided Wilcoxon signed-rank test showed significant improvement with LLM support [W(6) = 21.0, p = 0.0156, r = 0.879], with median accuracy increasing from 43.3% to 93.3% (median gain = 42.3%). Five of six students showed higher accuracy under the LLM-assisted condition, and no clear evidence of a sequence or carryover effect was detected (Mann-Whitney U = 7.0, p = 0.3758). Domain-level analyses indicated significant gains in all four categories (p < 0.05). Error-frequency analysis further showed marked reductions in high-frequency mistakes. The Likert-scale evaluation indicated high perceived clarity and coherence of LLM explanations, with favorable but more cautious instructor ratings. Discussion These pilot findings suggest that supervised LLM assistance may support art history question-answering and explanatory feedback. Future studies should validate these findings in larger cohorts, assess delayed learning retention, and examine open-ended, image-based, and higher-order art history tasks before curriculum-level implementation.

Yunting Zhang, Fan Zhang, Zi-Li Zhang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.