How consistent is the algorithm? Examining the intra- and inter-rater reliability of LLM-based writing assessment
Unlike constrained constructed-response assessments, the reliability of essay-based assessments has long been debated in language education research. Recently, AI—particularly large language models—has been proposed as a tool to enhance scoring reliability. This study examines the reliability of LLM-based scoring in wr...