ELBench is introduced, the first benchmark to evaluate all four requirements (General Capability, Safety and Trustworthiness, Basic Education, and High-Level Cultivation) on the same models under a common protocol, combining curated public sources with newly synthesized safety and cultivation data.
Abstract
Large language models are increasingly deployed in education as tutors, teaching assistants, and content generators. These roles place demands that ordinary question answering does not: a usable education-facing model is supposed to be accurate, safe under sensitive prompts, instructionally useful, and aligned with pedagogical goals at the same time. Existing benchmarks evaluate these requirements largely in isolation, so none assesses education-facing suitability as an integrated profile. We introduce ELBench, the first benchmark to evaluate all four requirements (General Capability, Safety and Trustworthiness, Basic Education, and High-Level Cultivation) on the same models under a common protocol, combining curated public sources with newly synthesized safety and cultivation data. We evaluate nine models, seven frontier general-purpose systems and two education-specialized variants, and report three findings. First, module-level profiles are more informative than a single aggregate: the top six models are statistically indistinguishable on overall score, yet their module leaders differ substantially, and safety is anti-correlated with practical teaching (r = -0.83). Second, the Chinese-developed models lead the safety module, the most discriminative in the suite; this advantage is largest on region-specific normative content and narrows, but does not vanish, on universal-harm content. Third, the two education-specialized models lead neither education module, and on High-Level Cultivation all models share a systematic blind spot: on the structured judgment task they converge on the same non-reference option, favoring pedagogical style over fit to the stated goal, so the module scores uniformly low and does not separate models. This raises, but does not resolve, whether domain post-training keeps pace with frontier systems on education tasks.
EduClaw-Bench is introduced, a benchmark that places an agent tutor in a continuous 30-day relationship with a simulated learner grounded in knowledge tracing (KT), whose knowledge-concept mastery, from a KT model trained on real-student data, drives its answers and is probed for learning gain across 55 scenarios.
Unggi Lee, Sookbun Lee, Yeil Jeong et al.· 0 citations
The modern classroom is an inherently multimodal environment, rich with the teacher's speech, student expressions, and interactive instructional resources. Effective AI assistants should therefore perceive this complex pedagogical process, not just answer text-based questions. However, existing evaluation methods are fundamentally misaligned. Current educational benchmarks primarily focus on unimodal and single-task evaluation, while existing multimodal benchmarks focus on content evaluation by testing MLLMs as students rather than the critical process evaluation by testing them as assistants. To fill this critical gap, we introduce Edu-Eval, the first large-scale benchmark focused on educational scenario-centric evaluation. Inspired by educational theories, we developed our Teacher-Student-Resource framework, which computationally operationalizes abstract pedagogical concepts into nine concrete evaluation tasks. Edu-Eval is constructed from over 3,000 real-world scenarios, comprising 75,900 annotated samples. Our evaluation of eight major MLLM series reveals a critical gap between current MLLM capabilities and the demands of real-world applications. Edu-Eval is more than a dataset. It is a diagnostic tool and a clear roadmap for future research, highlighting the urgent need for the community to shift its focus from isolated content evaluation to a holistic and scenario-centric understanding of the entire classroom ecosystem.
Zhiyi Duan, Jiangshan Guan, Qianli Xing· Proceedings of the 32nd ACM...· 0 citations
A comparative analysis of six LLMs for generating formative feedback on introductory Java programs containing predefined defects under controlled conditions reveals substantial cross-model variation, particularly in multi-defect scenarios.
Melina Najimi, Saba Yazdani, Marzieh Ahmadzadeh· Proceedings of the Canadian...· 0 citations
The results suggest that learner- and curriculum-aware alignment may matter more for effective tutoring than model category alone, and that such alignment is both measurable and improvable.
Benjamin Barlog, Hudson Craig, Ze-Dong Peng· IEEE International Conferenc...· 0 citations
EduMind is introduced, a unified tutoring and assessment platform designed around a dual-track evaluation model that demonstrates how assessment and tutoring can be unified into a seamless workflow, and remained operationally stable throughout all testing phases.
Dhyan Gowda, M. Aruna, P. Prasad et al.· International Journal of Sci...· 0 citations