Skip to content

Reliability, validity, and diagnostic evidence for multi-model LLM short-answer scoring

Sep 2026 · 0 citations · 1 references
Computer Science

TL;DR

Repeated multi-model OCG-PRES scoring provides reliability, validity, and diagnostic evidence for LLM-assisted short-answer scoring and supports cautious, evidence-based use as a scoring support tool rather than a replacement for human judgement.

Abstract

Large language models (LLMs) are increasingly used or proposed for educational scoring, but single-model and single-run evaluations provide limited evidence for assessment use. Short-answer scoring requires evidence about reliability, validity, severity, diagnostic value, and failure cases. This study evaluated repeated multi-model OCG-PRES guided LLM scoring for short-answer assessment. The analysis used 996 SciEntsBank responses. GPT, DeepSeek, and Qianwen each scored every response across three independent runs using five OCG-PRES dimensions: concept coverage, relation accuracy, reasoning completeness, contradiction control, and domain relevance. Scores were evaluated against official binary and five-category labels and compared with non-LLM baselines based on answer length, Jaccard keyword overlap, TF-IDF cosine similarity, and a combined traditional logistic model. Repeated-run reliability was high for all models, with ICC(3,k) = .977 for GPT, .992 for DeepSeek, and .981 for Qianwen. DeepSeek was the most stable across runs. GPT showed the strongest official-label alignment by AUC (.909), while Qianwen was stricter, with higher precision but lower recall under the fixed threshold = 3.0 rule. OCG-PRES scores followed expected diagnostic patterns across five official categories and outperformed all non-LLM baselines in AUC and F1. Repeated multi-model OCG-PRES scoring provides reliability, validity, and diagnostic evidence for LLM-assisted short-answer scoring. The findings support cautious, evidence-based use as a scoring support tool rather than a replacement for human judgement.

View source

Similar papers

Review Open access Sep 2026

Performance and Consistency of Large Language Models in Key Labor-Intensive Tasks of Systematic Reviews.

A suite of automated tools for automated batch processing that provide decision rationales and evidence enhances transparency and allows for human verification of AI decisions and provides a suite of automated tools for key SR tasks.

Yi-Ran Liu, Xi-Ling Wang, Zi-Xuan Zhou et al. · 0 citations
#natural language process... Preprint Sep 2026

When Consistency Does Not Mean Reliability: Evaluating Local LLM Judges Against Human Ratings

The results show that high self-consistency does not necessarily indicate high agreement with human judgments when using local LLMs as automatic judges, and highlight the need to evaluate both consistency and human alignment when using local LLMs as automatic judges.

Aakash Kumar Tiwari · 1 citation
Open access Sep 2026

A statistical framework for reliability evaluation of large language models

A statistical framework for quantitatively assessing LLM reliability under multiple evaluation criteria and shows that the proposed framework can identify reliability differences between different LLMs is reasonably robust to variations in indicator weights.

Yi Zhu · 0 citations
Review Open access 2026

Large Language Model-Based Automated Assessment: A Systematic Review, Taxonomy, and Implications for Personalized Learning

Research on large language model (LLM)-based automated assessment (AA) has expanded rapidly. Nevertheless, the literature remains fragmented across contributions, models, implementation configurations, datasets, and evaluation metrics, complicating efforts to identify approaches suitable for personalized learning. This...

H. D. Septama, A. E. Permanasari, R. Ferdiana · 0 citations
Conference Aug 2026

When Stable Answers are Not Enough: Evaluating Response Consistency and Medical Alignment in Consumer-Facing LLMs

Large Language Models (LLMs) are increasingly used by consumers as sources of health information. Evaluating the quality of such systems requires considering multiple quality dimensions rather than relying on a single indicator. Response consistency is often interpreted as a sign of reliability, but it remains unclear...

Line Praestegaard, Elda Paja · 0 citations
Review Open access Aug 2026

Large language models as judges: recent advances in LLM-based evaluation, critique, preference modeling, and feedback for text and code

This survey provides a comprehensive overview of recent advances in LLM-based evaluation, covering techniques, applications, and challenges across domains, with future directions emphasizing standardized protocols, uncertainty estimation, and human–AI collaboration.

M. Nadăş · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.