Dual-Signal Explainability for LLM-Based Automated Grading and Feedback: A Pilot Study
Abstract
The provision of prompt and customized feedback on student responses in descriptive questions continues to be a difficult problem for educators in virtual and hybrid education systems. The paper introduces the Explainable AI Grading Framework (XAIGF), which is a fusion of two paradigms - semantic evaluation through Large Language Models (LLMs) and deterministic lexical similarity analysis - to provide clear and pedagogically valuable feedback for automated long-answer scoring. XAIGF uses GPT-4o (temperature = 0.2) with a defined JSON output format, whereby LLM-generated rationales are combined with a Local Similarity Ratio (LSR) derived from Jaccard unigram overlap. In the pilot testing phase, the algorithm was evaluated against 20 student responses (120 responses graded) in a university-level Chemistry class (topic: Intermolecular Forces and Chemical Bonding; questions = 6; total marks = 26) with significant Pearson correlation $(\mathrm{r}$ 0.94; 95% confidence interval [CI]: 0.88-0.97) and Cohen's kappa coefficient $(\kappa=0.79)$ with human expert evaluations, as well as a massive cut in the time taken for grading (85%; 7.08× speedup). Perception surveys revealed substantial improvements in teachers' trust in feedback (+46 percentage points) and students' acceptance of generated feedback (+42 percentage points) after receiving the results from XAIGF. However, considering the limited scope and restricted domain of this pilot study, the outcomes obtained here must be considered preliminary indicators of potential generalizability. This framework proposes a hybrid approach to explainability in AI-enabled evaluation, integrating complementary semantic and lexical transparency cues under a responsible human-in the-loop deployment paradigm.