Skip to content
Review Open access

Evidence of impact and interpretational limits of generative AI in STEM education: a systematic review and meta-analysis on cognitive learning outcomes

Aug 2026 · Artificial Intelligence Review · Vol 59 · 3 citations · ⚡ 1 influential

TL;DR

The extent to which interactions with GAI enhance learning effectiveness and possible moderators, what challenges learners face when interacting with GAI systems, and which interventions support successful learner-GAI interaction are examined.

Abstract

This systematic review and meta-analysis examines the impact of generative artificial intelligence (GAI) on cognitive learning outcomes in STEM education. Prior research is growing but remains fragmented, often focusing on usability or single tools like ChatGPT rather than domain-specific cognitive effects. We therefore address these gaps by examining (1) the extent to which interactions with GAI enhance learning effectiveness and possible moderators, (2) what challenges learners face when interacting with GAI systems, and (3) which interventions support successful learner-GAI interaction. We meta-analyzed externally assessed cognitive outcomes (RQ1) and narratively synthesized reported learner challenges and supportive instructional interventions (RQ2-RQ3) when quantitative pooling was not feasible. Two pairs of raters independently screened and coded peer-reviewed quantitative studies published after 2017 that included a comparison/control group and examined cognitive learning in STEM involving learner-GAI interaction. A systematic search (ERIC, PsycINFO, Web of Science; updated May 7, 2026) and citation tracking yielded 85 eligible studies, of which 49 studies with 59 effect sizes met meta-analytic criteria. Most studies focused on higher education and used text-based GAI tools (e.g., ChatGPT). A random-effects meta-analysis shows an overall positive effects of GAI in STEM education, but the studies included exhibit a substantial heterogeneity (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$I^2=96.32\%$$\end{document}), and the prediction interval ranges from Hedge’s \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$g = -1.52$$\end{document} to \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$g = 3.20$$\end{document}. In line with this, funnel plot asymmetry suggests a potential publication bias. To account for potential publication bias, we also conducted a Robust Bayesian Meta-Analysis (RoBMA) and found that the overall positive effect of GAI in STEM education can be largely attributed to publication bias (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\mu =0.076\ \pm \ 0.254$$\end{document}), however a large heterogeneity remains (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\tau =1.190\ \pm \ 0.166$$\end{document}), which appears to be not associated to publication bias. To explain the heterogeneity, we conducted a moderator analysis and found that the learning outcome (knowledge vs. skills) as well the inversion-substitution-augmentation-redefinition (ISAR)-level, which compares the cognitive activity of the intervention and control in terms of the interactive-constructive-active-passive (ICAP)-level, both explain parts of the observed heterogeneity (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$R^2=12.5 \%$$\end{document} and \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$R^2=13.0 \%$$\end{document}, respectively). In contrast to the main effect, the moderator findings were robust under publication-bias correction using RoBMA. Apart from that, a coincidence analysis suggested that no combination of learning outcome type and ISAR level fulfills a sufficient condition for large effect sizes. Additionally, knowledge as a learning outcome type was found to be an almost, but not strictly, necessary condition for large effects when learning with generative AI. Overall, the RoBMA indicates that GAI appears promising in STEM if knowledge gains are targeted as learning outcomes, and when GAI is used to augment students’ learning activities, instead of substituting their activities, where it can be detrimental. Furthermore, numerous studies (N=33) reported large effects, but only because the cognitive activities of the students in the intervention and control groups were not comparable. Evidence for RQ2-RQ3 was limited and inconsistently reported; hence, these findings are presented as transparent, caveated qualitative insights rather than generalizable effect estimates. However, substantial unexplained heterogeneity and systematic underreporting of learner-level variables (AI literacy, metacognitive skills) and process-level mechanisms (task delegation, prompt quality, verification behaviors) indicate that the field needs to improve the theoretical models and has yet to measure the factors most likely to drive effectiveness. We propose six testable hypotheses and an integrative theoretical framework to guide future research toward understanding how, for whom, and under what conditions GAI supports STEM learning. The review protocol was preregistered on AsPredicted (ID:176450, URL: https://aspredicted.org/hpbd-dk75.pdf0).

Read PDF

Similar papers

Review Open access Aug 2026

Evaluating the Effectiveness of Flipped Lesson Models in Physical Education: A Systematic Review of Student Engagement and Learning Outcomes

Flipped learning relocates direct instruction to pre-class activities and uses class time for guided practice, feedback, and collaboration. Its effectiveness in physical education (PE), where learning spans cognitive, psychomotor, and social domains, remains unclear. This systematic review was registered in PROSPERO (CRD42024543676) and reported in accordance with PRISMA 2020. Five bibliographic databases and Google Scholar were searched for peer-reviewed intervention studies published from January 2018 to January 2024. Ten studies met the eligibility criteria, and their findings were synthesised thematically. Across secondary and tertiary PE settings, flipped lessons were generally associated with higher motivation, participation, self-efficacy, collaboration, conceptual understanding, and skill performance. Outcomes varied according to the quality and alignment of pre-class materials, students’ completion of preparatory work, teacher readiness, and contextual factors. Most interventions were short and geographically concentrated, and heterogeneity in study designs and measures precluded strong causal or long-term conclusions. Flipped lessons are a promising approach to active learning in PE, but benefits are not automatic. Effective implementation requires accessible, well-designed pre-class resources, purposeful in-class activities, teacher training, and support for students who struggle with self-directed preparation. Longer, more diverse, and methodologically robust studies are needed.

Ce Ren, Chun-Mei Li, R. Bailey et al. · 0 citations
Review Open access Sep 2026

Digital transformation in education: psychological perspectives on learning, engagement, and outcomes: a systematic review and meta-analysis

The process of digital transformation has developed new educational systems that use technology-based tools, including learning management systems, artificial intelligence, and gamified systems, as well as digital teaching methods. The studies show incomplete evidence about how this phenomenon affects learning engagement, motivation, cognitive processes, self-regulation skills, and institutional performance. This meta-analysis aimed to synthesize global evidence on the psychological perspectives of digital transformation in education, focusing on learning engagement, motivation, cognitive/metacognitive processes, self-regulation, and institutional factors. A systematic meta-analysis was conducted to assess 42 studies that employed a random-effects model with the inverse variance method as their analysis approach. The odds ratios to calculate effect sizes, which included 95% confidence intervals. The subgroup analyses were about five different psychological and institutional domains. The assessment of heterogeneity used I 2 statistics, while the publication bias assessment relied on funnel plots and Egger’s regression test. The results showed that digital transformation produced positive results in all areas which included learning engagement (OR = 1.79, 95% CI: 1.68–1.91), motivation, affective psychology (OR = 1.57, 95% CI: 1.21–2.03), cognitive and metacognitive psychology (OR = 1.79, 95% CI: 1.60–2.02) and self-regulation and learning psychology (OR = 1.66, 95% CI: 1.49–1.85) and institutional factors (OR = 1.71, 95% CI: 1.48–1.97). High motivation and self-regulation showed high domain heterogeneity, while the other domains maintained their consistency. The study showed no publication bias except for cognitive and metacognitive studies, which showed asymmetrical findings. Digital transformation in education brings about significant advantages for psychological and institutional results, which include better learning engagement and cognitive growth. The different factors that affect motivation and self-regulation show how contextual and pedagogical elements create an impact.

Ping-Ping Hou · 0 citations
Review Open access Aug 2026

A Scoping Review of Generative AI Usage and Human Cognition in Education and Professional Practice: Dual-Edged Cognitive Impacts and a Governance Framework

The Cognitive Impact Taxonomy (CIT) provides a structured approach for assessing Gen-AI’s influence on fluid and crystallized intelligence, key cognitive processes, and the trajectory of cognitive impacts leading to severity assessment, and the Integrated Cognitive Symbiosis Framework (ICSF) introduces a three-layer governance model to support responsible and cognitively healthy Gen-AI adoption.

K. Vidanage, M.J. Marapperuma, Shakya Dissanayake et al. · 0 citations
Review Open access Sep 2026

Harmonizing technology and pedagogy: A systematic review of Aidriven Competency learning models in education

Findings reveal that AIdriven interventions particularly those employing humancentered design, multimodal analytics, and outcomebased knowledge graph mapping—significantly improved competency gains, engagement, and instructional alignment.

Zahari Hamidon · 0 citations
Sep 2026

The Impact of Generative AI on Self-Efficacy: Evidence From a Meta-Analysis

Generative artificial intelligence (genAI) tools such as ChatGPT are increasingly used in education. While many studies have examined how genAI tools affect learning, we know less about how genAI tools influence motivational factors such as learners’ self-efficacy, despite the central role self-efficacy plays in motivation and learning. We conducted a meta-analysis of experimental and quasi-experimental research to examine the effect of genAI tools on learners’ self-efficacy. Studies published between November 2022 and February 2026 were systematically searched and screened, yielding 84 independent effects in our analysis. The results showed that using genAI during learning improved learners’ self-efficacy with a moderate effect size. However, the size of this effect depended on how genAI was used. Implementations where genAI acted as a tutor that guided students’ learning produced larger improvements in self-efficacy than uses where AI only (1) provided feedback, (2) answered questions, or (3) replaced learner work. Also, longer interventions produced stronger improvements than brief interactions. The improvements in self-efficacy did not differ across academic domains or learner age groups. These findings suggest that genAI can support learners’ self-efficacy when it is used to scaffold learning and encourage active engagement. The results highlight the importance of how genAI is integrated into instruction when designing AI-supported learning environments.

Hosain Heshmati, Ang Li, Jonathan G. Tullis · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.