Skip to content
Open access

Evaluating large language models for rubric-based essay grading in an undergraduate biology course

Jul 2026 · Journal of Microbiology & Biology Education · Vol 27 · 0 citations · 14 references
Medicine

Abstract

ABSTRACT This study examines how three large language models (LLMs), ChatGPT, Claude, and Gemini, assign grades to undergraduate-level essays in a biology course using a standardized rubric. Each LLM evaluated a data set of 200 essays under two prompting conditions: zero-shot (uncalibrated) and few-shot (calibrated using a small set of exemplar essays). LLM-assigned scores were directly compared with instructor-assigned scores, showing only moderate alignment with instructor grading, with variability observed across models and prompting strategies. Differences in grading behavior were also evident with different items on the rubric, with higher alignment for structural writing components and lower alignment for content- and reasoning-based criteria. Additionally, LLMs showed greater agreement with instructor scores than with one another, indicating substantial inter-model variability under identical grading conditions. These findings suggest that LLM grading outputs vary meaningfully across models, prompting strategies, and rubric components. In this context, LLMs may be best understood as tools that can support specific aspects of structured grading rather than as interchangeable evaluators.

Read PDF