Skip to content
Open access

A benchmark dataset with human validation for AI-assisted technical answer evaluation

Aug 2026 · Discover Education · Vol 5 · 0 citations · 26 references

TL;DR

This study presents DSA-RubricEval, a preliminarily reliability-assessed, rubric-based benchmark dataset intended to support exploratory research on pedagogically meaningful AI-assisted assessment of open-ended technical responses.

Abstract

Grading open-ended technical responses has been a longstanding issue in higher education. Despite advances in automated assessment, existing approaches often rely on holistic scoring, weakly validated annotations, and cognitive alignment, limiting pedagogical reliability and classroom adoption. To facilitate reliable and rubric-based automated evaluation in the Data Structures and Algorithms course, this study presents DSA-RubricEval, a pedagogically grounded and preliminarily reliability-assessed dataset. The dataset constitutes the primary contribution of this work, providing a structured benchmark aligned with Bloom’s taxonomy and validated through multi-rater Interclass Correlation Coefficient ICC analysis. The dataset consists of twelve expert-designed, Bloom-aligned, questions scored on multiple rubric dimensions using an ordinal scale. Five independent evaluators scored student responses, enabling rigorous validation of human judgement based on the Intraclass Correlation Coefficient (ICC) analysis. The results indicate good to excellent average-measure reliability across most rubric dimensions, justifying the use of aggregated human scores as aggregated reference labels for automated assessment. Automated scoring was explored as a proof of concept to demonstrate the applicability of the dataset for AI-assisted assessment and formulated as an ordinal, rubric-level prediction task and evaluated using pedagogically motivated agreement measures, showing high tolerance-based agreement with human judges. The proposed research identifies sources of assessor subjectivity and explores methods to mitigate them, while reducing the workload of grading as well as correlate with learning outcomes. Moreover, the proposed technique supports lower-order cognitive skills but is not best suited for higher-order cognitive tasks. Overall, this study introduces a preliminarily reliability-assessed, rubric-based benchmark dataset intended to support exploratory research on pedagogically meaningful AI-assisted assessment of open-ended technical responses.

Read PDF

Similar papers

Open access Aug 2026

EduFairBench: reproducible evaluation of large language models for educational assessment

E EduFairBench provides a reproducible methodology for jointly analyzing predictive performance, robustness, uncertainty, and feedback quality, providing a comprehensive methodological framework for the rigorous evaluation of LLM-based educational assessment systems.

W. Villegas-Ch., Aracely Mera-Navarrete, Fernando Zúñiga-Tello et al. · 0 citations
Conference Jul 2026

AutoQABench: A Three-Level UX Benchmark for Automated Evaluation of Open-Ended LLM Responses

Evaluating open-ended large language model responses remains difficult because response quality depends not only on factual correctness and task completion, but also on subjective and scenario-dependent user experience factors. Existing benchmarks and automatic evaluators are effective for coarse-grained assessment, yet often provide limited insight into how different types of quality failures affect user-perceived response quality. To address this gap, we propose AutoQABench, an initial three-level user experience benchmark for automated evaluation of open-ended LLM responses. The benchmark decomposes response quality into three progressively organized levels: basic acceptability constraints, scenario-specific task effectiveness, and preference-sensitive experiential quality. Based on this design, we construct a dataset covering representative scenarios, including summarization, elaboration, emotional support, and responses to misleading premises, together with expert-defined evaluation standards. We also develop an LLM-based modular evaluator that performs stepwise assessment and generates both intermediate judgments and a final rating. Experimental results show that AutoQABench achieves good agreement with expert judgments and supports analysis at the module, scenario, and final-grade levels. Rather than replacing human evaluation, AutoQABench is intended as a scalable auxiliary evaluator that helps identify overlooked risks, task-completion failures, and quality differences in open-ended LLM interactions.

Tianyou Wang, Fei Yuan, Chia-Ju Miao et al. · 0 citations
Conference Aug 2026

Can AI Grade Like Humans? Agreement and Reliability in Exam Assessment

Artificial intelligence (AI) is rapidly transforming higher education, with assessment emerging as one of its most impactful and contested applications. The automation of grading processes promises substantial improvements in efficiency, scalability, and standardisation. However, these potential benefits are accompanied by growing concerns about the reliability, consistency, fairness, and transparency of AI-based evaluation systems. These concerns are particularly critical in high-stakes academic assessment, where grading accuracy directly influences student outcomes and institutional credibility. Traditional grading relies on human evaluators who contribute contextual understanding, disciplinary expertise, and interpretative judgement. Nevertheless, human assessment is subject to limitations, including fatigue, variability, and potential bias. In contrast, AI systems are often assumed to provide more objective and consistent evaluations. Despite these assumptions, empirical evidence remains mixed. Recent studies suggest that while large language models (LLMs) can approximate human scoring patterns, they often exhibit weaker inter-rater agreement and inconsistent grading behaviour across different contexts. Moreover, existing research tends to focus on isolated performance indicators, such as accuracy or correlation, without simultaneously addressing both agreement between evaluators and internal consistency within evaluator groups. This gap limits the ability to fully evaluate the effectiveness and reliability of AI-based grading systems. JEL Codes: Keywords: artificial intelligence in education, automated assessment, inter-rater reliability, large language models, exam grading

R. Rodrigues, Andreia Ferreira, Nathalia Suchek · 0 citations
#small language model Open access Sep 2026

Human versus machine

For teachers to effectively use large-language-model(LLM)-based ratings in the formative or summative assessment of texts, it is essential to ensure that such ratings can assess student writing in a valid and reliable manner. This study investigates whether a validated human text-rating procedure (benchmark rating) can be replicated by an LLM-based rating procedure. We tested the replication with two genres of elementary school students’ text—narrative and instructive—using nine LLMs from three providers (OpenAI, Anthropic, Mistral). Each LLM generated three independent scores per text via structured, benchmark-aligned prompts that were then aggregated into a consensus score. Results showed that intrarater reliability was high to excellent, ICC(3, k) ≈ .68–.97, and alignment with human ratings ranged from moderate to strong, ICC(3, 1) ≈ .47–.85, with larger models consistently outperforming smaller ones. Systematic bias patterns emerged, varying by model and genre, indicating a need for calibration. Increasing output token windows and reducing temperature parameters mitigated truncation and schema-related failures. Although LLM-based benchmark ratings can approximate expert judgments and reduce the need for labor-intensive human triple coding, limitations remain regarding cost (for larger models), genre- and task specificity, and sensitivity to text presentation and student grade level—factors that constrain immediate classroom use, particularly for formative feedback.

Afra Sturm, Valentin Unger, Fabian Grünig · 0 citations
#artificial intelligence Review Sep 2026

A Human-in-the-Loop Framework for AI-Assisted Scoring in Large-Scale Writing Assessment

The integration of artificial intelligence (AI), particularly large language models (LLMs), into educational assessment has opened new opportunities to enhance the efficiency and scalability of grading processes. This study presents the design and validation of an AI-assisted scoring framework for written responses in a large-scale national assessment. The proposed approach focuses on short written texts of approximately 150-200 words and incorporates a human-in-the-loop strategy to preserve assessment quality while reducing manual workload. The study is grounded in a real operational context, using data from two recent editions of a nationwide test, each comprising approximately 5,000 student responses. We analyze the alignment between AI-generated scores and human raters across multiple rubric dimensions, as well as the impact of the proposed decision flow on pass/fail outcomes. Results show moderate to high agreement between the model and human evaluations in most dimensions, supporting the feasibility of AI assistance in this setting. Moreover, the proposed correction workflow identifies cases where human review is most valuable, enabling a more efficient allocation of expert effort. The findings suggest that AI-assisted scoring can be safely integrated into large-scale assessment processes only when combined with carefully designed human oversight. The paper concludes by discussing practical implications for deployment in national assessment systems and outlining future research directions, including longitudinal monitoring of model-human alignment and the analysis of potential cognitive bias introduced by AI-supported review workflows.

María Eugenia Curi, Germán Capdehourat, Isabel Amigo et al. · 0 citations
2026

How Prompt Design Shapes AI-Assisted Assessment: Reliability, Validity, and Learning Implications

This study examines the reliability and validity of generative AI in summative assessment, emphasizing how prompt design influences grading when applying a common rubric to complex student work. Fifteen business plans from a master’s-level course were evaluated by GPT-5 through Microsoft Copilot using three prompts of increasing rigor (basic, intermediate, rigorous). Each plan was scored in five independent runs per prompt, producing 225 AI evaluations. Analyses included intraclass correlation for consistency, severity contrasts, and convergence with instructor scores using correlation, error metrics, and Bland–Altman limits of agreement. Prompt design significantly shaped score distribution and strictness. The most rigorous prompt reduced inflated scores and aligned more closely with instructor judgments, yet it also underestimated performance and showed the greatest inconsistency. Single AI runs were unreliable, but averaging multiple evaluations improved stability. At the criterion level, AI struggled to match instructor ratings on commercial and economic viability, even under stricter prompts. Findings highlight the pedagogical implications of prompt sensitivity in AI-assisted grading. Reliability and fairness depend not only on rubric quality but also on evaluative instructions. Results support multi-run aggregation, bias-aware calibration, and hybrid human–AI models to ensure rigor and equity in technology-enhanced assessment. These findings inform the design of AI-enhanced assessment practices that support fair, transparent, and pedagogically aligned learning environments.

F. Miranda · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.