Skip to content

A Benchmark for LLM's Understanding of Middle School and High School Science Topics

Sep 2026 · 0 citations
Computer Science

Abstract

Large language models (LLMs) are increasingly integrated into educational settings, yet educators lack robust, standards-aligned tools to evaluate their effectiveness in K-12 science contexts. Existing benchmarks predominantly assess general language or advanced scientific reasoning, leaving a critical gap in understanding LLMs'performance on content directly relevant to secondary science curricula. To address this gap, we developed a comprehensive NGSS-aligned benchmark for both middle and high school science using a rigorous synthetic data pipeline, multi-judge validation, and item-level psychometric analysis. Nine open-weight LLMs were systematically evaluated using this benchmark, indicating that several smaller, locally deployable models achieved high accuracy across diverse science domains and question types. Our findings indicate that model size did not consistently predict performance, emphasizing the importance of intentional model selection for educational deployment. We then incorporated a human reviewer into the loop, reviewing the items generated by the LLMs for alignment with NGSS standards. The human review indicated that synthetically generated items were not in perfect alignment with the NGSS standards, indicating the benefits of human-in-the-loop item development, the need to explore the intersection of content and pedagogical knowledge, and the need to extend benchmarks to evaluate LLMs'capacity for interactive, evidence-based feedback in educational scenarios.

View source

Similar papers

Preprint Aug 2026

ELBench: A Multi-Dimensional Benchmark for Education-Facing Large Language Models

ELBench is introduced, the first benchmark to evaluate all four requirements (General Capability, Safety and Trustworthiness, Basic Education, and High-Level Cultivation) on the same models under a common protocol, combining curated public sources with newly synthesized safety and cultivation data.

Yi-Lin Jiang, Xiao-Rong Zhu, Fei Tan et al. · 0 citations
Open access Aug 2026

Comparative Evaluation of Large Language Models in Computer Programming Education

A comparative analysis of six LLMs for generating formative feedback on introductory Java programs containing predefined defects under controlled conditions reveals substantial cross-model variation, particularly in multi-defect scenarios.

Melina Najimi, Saba Yazdani, Marzieh Ahmadzadeh · 0 citations
#human-computer interacti... Preprint Oct 2026

Understanding Student Use of Large Language Models Across Computer Science Subfields

This research full paper examines how undergraduate students use large language models (LLMs) across computer science subfields. As LLMs become increasingly integrated into computing education, understanding how their use varies across technical and pedagogical contexts is essential for designing effective, subfield-aw...

S. Nizamani, Yoonjeong Lee, Nikitha Donekal Chandrashekar et al. · 0 citations
Open access Sep 2026

ENHANCING FINAL-YEAR STUDENTS' ENGLISH PROFICIENCY THROUGH STRUCTURED TOEFL PREPARATION TRAINING

The global demand for English proficiency has made standardized assessments such as the TOEFL critical benchmarks for academic and professional success. In alignment with Sustainable Development Goal 4 (Quality Education), this study evaluates the effectiveness of a structured TOEFL Preparation training model for 88 fi...

Ida Zuraida, H. Hendar, Heri Heryono et al. · 0 citations
Open access 2026

EduDesignEval: A Structured Evaluation Framework for Educational Program Designer Using Large Language Models

The rapid advancement of Large Language Models (LLMs) has created unprecedented opportunities for transforming educational program design. However, the lack of standardized evaluation frameworks raises concerns regarding the content quality and pedagogical effectiveness. This study proposes EduDesignEval, a comprehensi...

Ayanbek Serikov, A. Biloshchytskyi, A. Mukhatayev et al. · 0 citations
#small language model Open access Sep 2026

Human versus machine

Although LLM-based benchmark ratings can approximate expert judgments and reduce the need for labor-intensive human triple coding, limitations remain regarding cost, genre- and task specificity, and sensitivity to text presentation and student grade level—factors that constrain immediate classroom use, particularly for...

Afra Sturm, Valentin Unger, Fabian Grünig · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 30, 2026

This game-playing AI is the new champ at Stratego

Able to defeat top-ranked human players and more efficient than other models, the new system could help decision-makers in military maneuvers or business negotiations.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.