Large language models (LLMs) are increasingly integrated into educational settings, yet educators lack robust, standards-aligned tools to evaluate their effectiveness in K-12 science contexts. Existing benchmarks predominantly assess general language or advanced scientific reasoning, leaving a critical gap in understanding LLMs'performance on content directly relevant to secondary science curricula. To address this gap, we developed a comprehensive NGSS-aligned benchmark for both middle and high school science using a rigorous synthetic data pipeline, multi-judge validation, and item-level psychometric analysis. Nine open-weight LLMs were systematically evaluated using this benchmark, indicating that several smaller, locally deployable models achieved high accuracy across diverse science domains and question types. Our findings indicate that model size did not consistently predict performance, emphasizing the importance of intentional model selection for educational deployment. We then incorporated a human reviewer into the loop, reviewing the items generated by the LLMs for alignment with NGSS standards. The human review indicated that synthetically generated items were not in perfect alignment with the NGSS standards, indicating the benefits of human-in-the-loop item development, the need to explore the intersection of content and pedagogical knowledge, and the need to extend benchmarks to evaluate LLMs'capacity for interactive, evidence-based feedback in educational scenarios.
ELBench is introduced, the first benchmark to evaluate all four requirements (General Capability, Safety and Trustworthiness, Basic Education, and High-Level Cultivation) on the same models under a common protocol, combining curated public sources with newly synthesized safety and cultivation data.
Yi-Lin Jiang, Xiao-Rong Zhu, Fei Tan et al.· 0 citations
A comparative analysis of six LLMs for generating formative feedback on introductory Java programs containing predefined defects under controlled conditions reveals substantial cross-model variation, particularly in multi-defect scenarios.
Melina Najimi, Saba Yazdani, Marzieh Ahmadzadeh· Proceedings of the Canadian...· 0 citations
This research full paper examines how undergraduate students use large language models (LLMs) across computer science subfields. As LLMs become increasingly integrated into computing education, understanding how their use varies across technical and pedagogical contexts is essential for designing effective, subfield-aw...
S. Nizamani, Yoonjeong Lee, Nikitha Donekal Chandrashekar et al.· 0 citations
The global demand for English proficiency has made standardized assessments such as the TOEFL critical benchmarks for academic and professional success. In alignment with Sustainable Development Goal 4 (Quality Education), this study evaluates the effectiveness of a structured TOEFL Preparation training model for 88 fi...
Ida Zuraida, H. Hendar, Heri Heryono et al.· English Journal Literacy UTa...· 0 citations
The rapid advancement of Large Language Models (LLMs) has created unprecedented opportunities for transforming educational program design. However, the lack of standardized evaluation frameworks raises concerns regarding the content quality and pedagogical effectiveness. This study proposes EduDesignEval, a comprehensi...
Ayanbek Serikov, A. Biloshchytskyi, A. Mukhatayev et al.· International Journal of Inf...· 0 citations
Although LLM-based benchmark ratings can approximate expert judgments and reduce the need for labor-intensive human triple coding, limitations remain regarding cost, genre- and task specificity, and sensitivity to text presentation and student grade level—factors that constrain immediate classroom use, particularly for...
Able to defeat top-ranked human players and more efficient than other models, the new system could help decision-makers in military maneuvers or business negotiations.