Skip to content
Open access

A RAG-Enhanced Human-in-the-Loop Framework for Automated Assessment of Engineering Laboratory Reports

2026 · International Journal of Advanced Computer Science and Applications · 0 citations · 25 references

TL;DR

The results demonstrate that the integration of Retrieval-Augmented Generation, rubric-based evaluation, and Human-in-the-Loop validation constitutes an effective approach for AI-supported assessment of engineering laboratory reports, maintaining instructor oversight and educational integrity.

Abstract

Recent developments in Large Language Models (LLMs) have created new opportunities to automate educational assessment and reduce workload for instructors. However, concerns regarding grading consistency, transparency, and pedagogical reliability continue to limit the adoption of fully automated assessment systems. In this study, we propose a Human-in-the-Loop framework for the automated evaluation of engineering laboratory reports based on Retrieval-Augmented Generation (RAG). The proposed framework integrates text extraction, structured content extraction, contextual retrieval from grading rubrics and laboratory resources, rubric-based evaluation, automatic feedback generation, and instructor validation into a unified grading workflow. The RAG module retrieves contextual information so that the language model can generate assessments that align with the course goals and are based on educational information about the specific assignments, thus ensuring consistent evaluation based on the rubric. The system was tested using a set of 56 laboratory reports collected from undergraduate courses in Electrical and Electronics Engineering. The experimental results indicate a strong agreement between the grades provided by the AI and the instructor, with a Pearson Correlation Coefficient of 0.988, a Mean Absolute Error (MAE) of 3.55, and a Root Mean Squared Error (RMSE) of 3.77. Besides, 89.29% of the reports were scored within ±5 points of the instructor scores. The grading time was reduced from 392 minutes to 84 minutes, a workload reduction of 78.57%. The results demonstrate that the integration of Retrieval-Augmented Generation, rubric-based evaluation, and Human-in-the-Loop validation constitutes an effective approach for AI-supported assessment of engineering laboratory reports, maintaining instructor oversight and educational integrity.

Read PDF

Similar papers

Review Open access Aug 2026

Human-in-the-Loop LLM Assessment for Programming Education: Design and Empirical Validation

Programming instructors face the challenge of providing prompt and consistent feedback, yet manual grading becomes unsustainable in large classes. While Large Language Models (LLMs) offer grading assistance, most research is based on English-language contexts and offline assessments, creating uncertainty about their dependability and the extent of human oversight needed. This study aimed to create and assess an evaluation ecosystem that integrates learning management, AI-driven task creation, LLM grading, and human review for programming courses taught in Indonesian. Employing an ADDIE-based Research and Development approach, the system was implemented for 109 students. For grading validation, instructors independently evaluated 50 assignments without access to AI predictions. The agreement was substantial, with a mean absolute error (MAE) of 4.14, a Pearson correlation of 0.986 within a 95 percent confidence interval ranging from 0.975 to 0.992, and an intraclass correlation coefficient (ICC) of 0.977. Instructor adjustments were more frequent for open-ended tasks (34.9 percent) compared to quizzes (20.2 percent). These results contributed to the development of the task-dependent human calibration (TDHC) model. The system attained a System Usability Scale (SUS) score of 88.5 and cut grading time by 87.5 percent, facilitating focused instructor review in LLM-supported programming assessments.

A. Ibrahim, Runal Rezkiawan · 0 citations
#artificial intelligence Review Sep 2026

A Human-in-the-Loop Framework for AI-Assisted Scoring in Large-Scale Writing Assessment

The integration of artificial intelligence (AI), particularly large language models (LLMs), into educational assessment has opened new opportunities to enhance the efficiency and scalability of grading processes. This study presents the design and validation of an AI-assisted scoring framework for written responses in a large-scale national assessment. The proposed approach focuses on short written texts of approximately 150-200 words and incorporates a human-in-the-loop strategy to preserve assessment quality while reducing manual workload. The study is grounded in a real operational context, using data from two recent editions of a nationwide test, each comprising approximately 5,000 student responses. We analyze the alignment between AI-generated scores and human raters across multiple rubric dimensions, as well as the impact of the proposed decision flow on pass/fail outcomes. Results show moderate to high agreement between the model and human evaluations in most dimensions, supporting the feasibility of AI assistance in this setting. Moreover, the proposed correction workflow identifies cases where human review is most valuable, enabling a more efficient allocation of expert effort. The findings suggest that AI-assisted scoring can be safely integrated into large-scale assessment processes only when combined with carefully designed human oversight. The paper concludes by discussing practical implications for deployment in national assessment systems and outlining future research directions, including longitudinal monitoring of model-human alignment and the analysis of potential cognitive bias introduced by AI-supported review workflows.

María Eugenia Curi, Germán Capdehourat, Isabel Amigo et al. · 0 citations
Open access Aug 2026

Lightweight Llama Models with Experts’ Curated RAG for Electrical Engineering Education: An Exploratory Comparison

The results show that the proposed RAG agent substantially improves the lightweight base model and produces transparent, syllabus-grounded answers that experts rated as correct and concise, while GPT-4.5 retains an edge on longer, multistep, and topology-intensive tasks.

André Rocha, Paulo Oliveira, João Ferreira et al. · 0 citations
Open access 2026

Leveraging large language models for scalable analysis of the end-of-the-course student feedback

Analyzing open-ended student feedback in course evaluations is a laborintensive task due to the unstructured and complex nature of natural language. While Large Language Models (LLMs) offer significant potential for automation, a welldefined methodology for their application in analyzing student feedback remains underdeveloped. This paper addresses this gap by proposing an LLM-based feedback analytics pipeline designed to transform students’ open-ended feedback into structured, actionable insights. The pipeline consists of three sequential stages: (i) segmenting student feedback into semantic units and assigning polarity (sentiment) to those units; (ii) topical classification of semantic units, and (iii) summarization of units within each topical category. By systematizing these processes, the proposed method enables educators and course managers to efficiently derive meaningful patterns from vast datasets of student opinions. We evaluated the proposed method using a comprehensive dataset from several editions of a U.S. university course, yielding encouraging results of the method’s effectiveness. This research provides a scalable, generic methodology for (semi-)automated feedback analysis, ultimately supporting data-informed improvements in teaching and course management.

J. Jovanović, Irena Vodenska, V. Devedžić · 0 citations
Conference Jul 2026

SpartanAI: An AI Agent for Automated Course Review

Quality assurance in online and blended course design remains labor-intensive, costly, and difficult to scale across institutions. Faculty often lack timely access to instructional design support aligned with established quality standards such as the Quality Matters (QM) rubric. While instructional designers bring deep expertise in online course design, faculty often lack equivalent familiarity with evidence-based design principles; SpartanAI provides these instructors with structured, rubricgrounded feedback that would otherwise require specialized expertise or a formal review process to obtain. While recent advances in large language models (LLMs) have enabled automated analysis of textual artifacts, most AI applications in education remain student-facing and do not provide systematic, rubricaligned course review services for instructors. In this paper we present SpartanAI, an LLM-enabled AI Framework that automates QM-aligned course review through structured learning management system (LMS) ingestion, module-level content extraction, semantic indexing, and rubric-by-rubric evaluation. Instructors upload exported course packages which are parsed into structured representations, indexed using vector embeddings, and evaluated rubric by rubric. For each rubric, the system retrieves relevant course artifacts, generates alignment scores, identifies supporting evidence, and provides actionable revision guidance. Initial evaluation against QM-certified reviewer feedback on a pilot set of courses shows promising alignment in rubric-level scoring and feedback relevance. These results suggest that structured LLM-based evaluation pipelines can augment institutional course quality assurance processes and provide scalable, instructor-centered AI support.

Pratik Korat, Darshan Patel, M. Eirinaki et al. · 0 citations
Open access Aug 2026

A GenAI-Based Adaptive Tutoring ana Intelligent Assessment Framework for Personalized Learning

EduMind is introduced, a unified tutoring and assessment platform designed around a dual-track evaluation model that demonstrates how assessment and tutoring can be unified into a seamless workflow, and remained operationally stable throughout all testing phases.

Dhyan Gowda, M. Aruna, P. Prasad et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.