Skip to content
Review Open access

Evidence of impact and interpretational limits of generative AI in STEM education: a systematic review and meta-analysis on cognitive learning outcomes

Aug 2026 · Artificial Intelligence Review · 2 citations

Abstract

This systematic review and meta-analysis examines the impact of generative artificial intelligence (GAI) on cognitive learning outcomes in STEM education. Prior research is growing but remains fragmented, often focusing on usability or single tools like ChatGPT rather than domain-specific cognitive effects. We therefore address these gaps by examining (1) the extent to which interactions with GAI enhance learning effectiveness and possible moderators, (2) what challenges learners face when interacting with GAI systems, and (3) which interventions support successful learner-GAI interaction. We meta-analyzed externally assessed cognitive outcomes (RQ1) and narratively synthesized reported learner challenges and supportive instructional interventions (RQ2-RQ3) when quantitative pooling was not feasible. Two pairs of raters independently screened and coded peer-reviewed quantitative studies published after 2017 that included a comparison/control group and examined cognitive learning in STEM involving learner-GAI interaction. A systematic search (ERIC, PsycINFO, Web of Science; updated May 7, 2026) and citation tracking yielded 85 eligible studies, of which 49 studies with 59 effect sizes met meta-analytic criteria. Most studies focused on higher education and used text-based GAI tools (e.g., ChatGPT). A random-effects meta-analysis shows an overall positive effects of GAI in STEM education, but the studies included exhibit a substantial heterogeneity ( $$I^2=96.32\%$$ ), and the prediction interval ranges from Hedge’s $$g = -1.52$$ to $$g = 3.20$$ . In line with this, funnel plot asymmetry suggests a potential publication bias. To account for potential publication bias, we also conducted a Robust Bayesian Meta-Analysis (RoBMA) and found that the overall positive effect of GAI in STEM education can be largely attributed to publication bias ( $$\mu =0.076\ \pm \ 0.254$$ ), however a large heterogeneity remains ( $$\tau =1.190\ \pm \ 0.166$$ ), which appears to be not associated to publication bias. To explain the heterogeneity, we conducted a moderator analysis and found that the learning outcome (knowledge vs. skills) as well the inversion-substitution-augmentation-redefinition (ISAR)-level, which compares the cognitive activity of the intervention and control in terms of the interactive-constructive-active-passive (ICAP)-level, both explain parts of the observed heterogeneity ( $$R^2=12.5 \%$$ and $$R^2=13.0 \%$$ , respectively). In contrast to the main effect, the moderator findings were robust under publication-bias correction using RoBMA. Apart from that, a coincidence analysis suggested that no combination of learning outcome type and ISAR level fulfills a sufficient condition for large effect sizes. Additionally, knowledge as a learning outcome type was found to be an almost, but not strictly, necessary condition for large effects when learning with generative AI. Overall, the RoBMA indicates that GAI appears promising in STEM if knowledge gains are targeted as learning outcomes, and when GAI is used to augment students’ learning activities, instead of substituting their activities, where it can be detrimental. Furthermore, numerous studies ( N =33) reported large effects, but only because the cognitive activities of the students in the intervention and control groups were not comparable. Evidence for RQ2-RQ3 was limited and inconsistently reported; hence, these findings are presented as transparent, caveated qualitative insights rather than generalizable effect estimates. However, substantial unexplained heterogeneity and systematic underreporting of learner-level variables (AI literacy, metacognitive skills) and process-level mechanisms (task delegation, prompt quality, verification behaviors) indicate that the field needs to improve the theoretical models and has yet to measure the factors most likely to drive effectiveness. We propose six testable hypotheses and an integrative theoretical framework to guide future research toward understanding how, for whom, and under what conditions GAI supports STEM learning. The review protocol was preregistered on AsPredicted (ID:176450, URL: https://aspredicted.org/hpbd-dk75.pdf0 ).

Read PDF