Skip to content

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Jul 2026

Evaluating large language models for rubric-based essay grading in an undergraduate biology course

ABSTRACT This study examines how three large language models (LLMs), ChatGPT, Claude, and Gemini, assign grades to undergraduate-level essays in a biology course using a standardized rubric. Each LLM evaluated a data set of 200 essays under two prompting conditions: zero-shot (uncalibrated) and few-shot (calibrated using a small set of exemplar essays). LLM-assigned scores were directly compared with instructor-assigned scores, showing only moderate alignment with instructor grading, with variability observed across models and prompting strategies. Differences in grading behavior were also evident with different items on the rubric, with higher alignment for structural writing components and lower alignment for content- and reasoning-based criteria. Additionally, LLMs showed greater agreement with instructor scores than with one another, indicating substantial inter-model variability under identical grading conditions. These findings suggest that LLM grading outputs vary meaningfully across models, prompting strategies, and rubric components. In this context, LLMs may be best understood as tools that can support specific aspects of structured grading rather than as interchangeable evaluators.

M. Naidu, Nikolas S. Montaquila, Jessica P Roa et al. · 0 citations