Skip to content

Author

M. Craciun

6 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Aug 2026

CreditTrace-LLM: Auditing Rubric-Point Responsiveness and Credit Locality in LLM-Based Automated Grading

Large language models are increasingly used for automated grading, but final-score agreement does not reveal whether credit is assigned to the rubric element affected by a response change. This study introduces CreditTrace-LLM, a controlled framework for auditing directional responsiveness, credit locality, paraphrase stability, and correspondence with expert evaluations in rubric-guided grading. The evaluation used 500 response families across ten questions, equally divided between technical and argumentative tasks. Each family included a baseline response (C0), a positive intervention adding rubric-relevant evidence (C+), a negative intervention removing or weakening such evidence (C−), and a placebo paraphrase (CP) constructed through lexical or syntactic reformulation with the intention of preserving rubric-relevant meaning. Eight local open-weight LLM graders produced 15,999 structurally valid Gold-point-level outputs from 16,000 expected evaluations. Targeted Gold-point scores changed in the expected direction in 75.7% of C+ comparisons and 50.2% of C− comparisons. Localized directional success was lower under C+ and C−, at 31.5% and 21.3%, respectively, indicating frequent non-target score changes. Under CP, total-score stability was 75.4%, while complete-profile stability was 73.6%. Direction concordance with the mean expert score change was 69.8% for C+, 47.4% for C−, and 18.2% for CP. The CP condition also showed substantial variability in expert scoring, particularly for argumentative responses, indicating that intended rubric-relevant meaning preservation did not guarantee score invariance. Final-score behavior alone is insufficient for validating rubric-guided LLM grading. CreditTrace-LLM therefore evaluates target responsiveness, credit locality, paraphrase stability, and traceable Gold-point-level outputs as complementary diagnostic dimensions rather than as predefined criteria for acceptable grading performance.

Cătălin Anghel, A. Anghel, M. Craciun et al. · 0 citations
#small language model Open access Sep 2026

RubricAdapt-LLM: Measuring Criterion-Level Adaptation and Score Shifts Under Alternative Grading Rubrics

Background: Large language models are increasingly used for rubric-based grading, but it remains unclear whether they adapt selectively when criterion weights change while the response, criterion definitions, and total score remain fixed. This study examined whether alternative point allocations produce targeted criterion-level adaptation or broader grading instability. Methods: A controlled paired design was applied to 1000 student responses, comprising 500 technical and 500 argumentative answers. Eight local open-weight LLMs evaluated every response under two analytic rubrics totaling 10 points. One point was transferred from Clarity to Completeness for technical responses and from Clarity to Dialecticality for argumentative responses. Of 16,000 expected evaluations, 15,999 were structurally valid, yielding 7999 complete cross-rubric pairs. Model outputs were analyzed for total-score shifts, affected-criterion adaptation, stability of unaffected criteria, model–human alignment, and correspondence with differences between two human evaluation conditions. Results: All models assigned lower mean scores under Rubric B, with mean shifts ranging from −1.239 to −0.182 points. Adaptation mechanisms differed substantially across models and response types. gemma3:4b frequently preserved technical total scores through compensating criterion changes, whereas llama3.1:8b showed extensive spillover into unaffected criteria. The Qwen models generally produced smaller total-score reductions and greater stability in unchanged dimensions. The mean score was 0.637 points higher under the human Rubric B condition than under the human Rubric A condition, although the two conditions were applied by different evaluator pairs; every model shifted negatively, and the lowest overall shift error was obtained by qwen3:4b at 1.498 points. Conclusions: Rubric sensitivity did not consistently imply localized or criterion-consistent adaptation. Reliable evaluation of LLM graders therefore requires separate analysis of total scores, affected criteria, unaffected criteria, human alignment, and cross-condition shift correspondence.

Cătălin Anghel, A. Anghel, Adina Cocu et al. · 0 citations
Open access Aug 2026

From Medical Records to AI-Ready Datasets: A Practical Guide for Clinical Researchers

A physician-facing Clinical AI-Readiness Guide for preparing medical datasets before AI-based analysis to improve collaboration between clinical and technical teams and reduce preventable dataset-related failures in medical AI research is proposed.

Cătălin Anghel, A. Anghel, M. Craciun et al. · 0 citations
Open access Aug 2026

ConsensusGrade: A Human-Variability-Aware Framework for Evaluating LLM-Based Automated Grading

ConsensusGrade provides a diagnostic framework for interpreting automated scores relative to observed human variability; inside-envelope rates should not be interpreted as stand-alone measures of grading accuracy.

Cătălin Anghel, A. Anghel, Mihai Vlase et al. · 0 citations
Open access Aug 2026

GradeDrift-LLM: Measuring Student-History-Induced Score Drift in LLM-Based Automated Grading

Student-history metadata can influence LLM-generated grading scores despite explicit instructions to ignore it, and future LLM-based grading systems should separate answer-based scoring from learner-context-based personalization and validate score invariance under controlled learner-context variations.

Cătălin Anghel, A. Anghel, M. Craciun et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.