Large language models are increasingly used for automated grading, but final-score agreement does not reveal whether credit is assigned to the rubric element affected by a response change. This study introduces CreditTrace-LLM, a controlled framework for auditing directional responsiveness, credit locality, paraphrase stability, and correspondence with expert evaluations in rubric-guided grading. The evaluation used 500 response families across ten questions, equally divided between technical and argumentative tasks. Each family included a baseline response (C0), a positive intervention adding rubric-relevant evidence (C+), a negative intervention removing or weakening such evidence (C−), and a placebo paraphrase (CP) constructed through lexical or syntactic reformulation with the intention of preserving rubric-relevant meaning. Eight local open-weight LLM graders produced 15,999 structurally valid Gold-point-level outputs from 16,000 expected evaluations. Targeted Gold-point scores changed in the expected direction in 75.7% of C+ comparisons and 50.2% of C− comparisons. Localized directional success was lower under C+ and C−, at 31.5% and 21.3%, respectively, indicating frequent non-target score changes. Under CP, total-score stability was 75.4%, while complete-profile stability was 73.6%. Direction concordance with the mean expert score change was 69.8% for C+, 47.4% for C−, and 18.2% for CP. The CP condition also showed substantial variability in expert scoring, particularly for argumentative responses, indicating that intended rubric-relevant meaning preservation did not guarantee score invariance. Final-score behavior alone is insufficient for validating rubric-guided LLM grading. CreditTrace-LLM therefore evaluates target responsiveness, credit locality, paraphrase stability, and traceable Gold-point-level outputs as complementary diagnostic dimensions rather than as predefined criteria for acceptable grading performance.
Cătălin Anghel, A. Anghel, M. Craciun et al.· Computers· 0 citations
A physician-facing Clinical AI-Readiness Guide for preparing medical datasets before AI-based analysis to improve collaboration between clinical and technical teams and reduce preventable dataset-related failures in medical AI research is proposed.
Cătălin Anghel, A. Anghel, M. Craciun et al.· Journal of Clinical Medicine· 0 citations
ConsensusGrade provides a diagnostic framework for interpreting automated scores relative to observed human variability; inside-envelope rates should not be interpreted as stand-alone measures of grading accuracy.
Cătălin Anghel, A. Anghel, Mihai Vlase et al.· Applied System Innovation· 0 citations
Student-history metadata can influence LLM-generated grading scores despite explicit instructions to ignore it, and future LLM-based grading systems should separate answer-based scoring from learner-context-based personalization and validate score invariance under controlled learner-context variations.
Cătălin Anghel, A. Anghel, M. Craciun et al.· Machine Learning and Knowled...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.