Psychological empowerment and relational catalysis: a quasi-experimental study on graduate education evaluation reform in the AI era
Abstract
Introduction The rise of generative AI (e.g., ChatGPT, DeepSeek) challenges traditional graduate evaluation systems centered on individual theses and standardized examinations. Current institutional responses—task redesign and AI skills training—have shown limited effectiveness, as they overlook the psychological internalization process, teacher-student relational dynamics, and their synergistic effects. Methods This study conducted a three-arm quasi-experiment with 155 doctoral students in education over one semester. Group A (traditional model) received no AI guidance; Group B received AI skills training only; Group C received comprehensive reform combining AI training, practical team-based projects, and humble mentor leadership training. Psychological empowerment was measured as a mediator, and humble mentor leadership as a moderator. Data were analyzed using regression with cluster-robust standard errors and bootstrap tests. Results Group C scored significantly higher than Groups A and B on course satisfaction, professional competence, and psychological empowerment (all p < 0.001). Psychological empowerment showed a significant indirect association with learning outcomes (indirect effect = 0.22, 95% CI [0.12, 0.34]), with meaning and autonomy as the core dimensions. Humble mentor leadership significantly moderated the engagement–empowerment path within Group C (β = 0.24, p < 0.001), with the conditional indirect effect significant only under high humble leadership. Discussion These findings suggest that the synergy of practical task redesign, psychological empowerment, and humble mentor leadership may be associated with positive student outcomes in AI-era educational reform. However, due to the cluster-assigned design with one cluster per condition and non-equivalent assessment tasks across groups, all results should be interpreted as exploratory and hypothesis-generating rather than causal. Future research should employ multi-cluster or individually randomized designs with common assessment tasks to permit valid causal inference.