Skip to content

Author

Jochen Kuhn

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Review Open access Aug 2026

Evidence of impact and interpretational limits of generative AI in STEM education: a systematic review and meta-analysis on cognitive learning outcomes

This systematic review and meta-analysis examines the impact of generative artificial intelligence (GAI) on cognitive learning outcomes in STEM education. Prior research is growing but remains fragmented, often focusing on usability or single tools like ChatGPT rather than domain-specific cognitive effects. We therefore address these gaps by examining (1) the extent to which interactions with GAI enhance learning effectiveness and possible moderators, (2) what challenges learners face when interacting with GAI systems, and (3) which interventions support successful learner-GAI interaction. We meta-analyzed externally assessed cognitive outcomes (RQ1) and narratively synthesized reported learner challenges and supportive instructional interventions (RQ2-RQ3) when quantitative pooling was not feasible. Two pairs of raters independently screened and coded peer-reviewed quantitative studies published after 2017 that included a comparison/control group and examined cognitive learning in STEM involving learner-GAI interaction. A systematic search (ERIC, PsycINFO, Web of Science; updated May 7, 2026) and citation tracking yielded 85 eligible studies, of which 49 studies with 59 effect sizes met meta-analytic criteria. Most studies focused on higher education and used text-based GAI tools (e.g., ChatGPT). A random-effects meta-analysis shows an overall positive effects of GAI in STEM education, but the studies included exhibit a substantial heterogeneity ( $$I^2=96.32\%$$ ), and the prediction interval ranges from Hedge’s $$g = -1.52$$ to $$g = 3.20$$ . In line with this, funnel plot asymmetry suggests a potential publication bias. To account for potential publication bias, we also conducted a Robust Bayesian Meta-Analysis (RoBMA) and found that the overall positive effect of GAI in STEM education can be largely attributed to publication bias ( $$\mu =0.076\ \pm \ 0.254$$ ), however a large heterogeneity remains ( $$\tau =1.190\ \pm \ 0.166$$ ), which appears to be not associated to publication bias. To explain the heterogeneity, we conducted a moderator analysis and found that the learning outcome (knowledge vs. skills) as well the inversion-substitution-augmentation-redefinition (ISAR)-level, which compares the cognitive activity of the intervention and control in terms of the interactive-constructive-active-passive (ICAP)-level, both explain parts of the observed heterogeneity ( $$R^2=12.5 \%$$ and $$R^2=13.0 \%$$ , respectively). In contrast to the main effect, the moderator findings were robust under publication-bias correction using RoBMA. Apart from that, a coincidence analysis suggested that no combination of learning outcome type and ISAR level fulfills a sufficient condition for large effect sizes. Additionally, knowledge as a learning outcome type was found to be an almost, but not strictly, necessary condition for large effects when learning with generative AI. Overall, the RoBMA indicates that GAI appears promising in STEM if knowledge gains are targeted as learning outcomes, and when GAI is used to augment students’ learning activities, instead of substituting their activities, where it can be detrimental. Furthermore, numerous studies ( N =33) reported large effects, but only because the cognitive activities of the students in the intervention and control groups were not comparable. Evidence for RQ2-RQ3 was limited and inconsistently reported; hence, these findings are presented as transparent, caveated qualitative insights rather than generalizable effect estimates. However, substantial unexplained heterogeneity and systematic underreporting of learner-level variables (AI literacy, metacognitive skills) and process-level mechanisms (task delegation, prompt quality, verification behaviors) indicate that the field needs to improve the theoretical models and has yet to measure the factors most likely to drive effectiveness. We propose six testable hypotheses and an integrative theoretical framework to guide future research toward understanding how, for whom, and under what conditions GAI supports STEM learning. The review protocol was preregistered on AsPredicted (ID:176450, URL: https://aspredicted.org/hpbd-dk75.pdf0 ).

Chiara Boolzen, Jochen Kuhn, Salome Flegr et al. · 2 citations
Preprint Aug 2026

AI-Augmented Inquiry and Regulation in Hybrid Systems: A Control Allocation Architecture for Preserving Epistemic Agency in Hybrid Human-AI Cognition

Generative artificial intelligence (genAI) systems are increasingly integral to epistemic processes such as hypothesis generation, explanation construction, and decision-making. Although they reliably enhance performance, emerging evidence reveals a metacognitive dilemma: as external generative capacity increases, internal monitoring, calibration, and cognitive engagement may decline. This reflects a redistribution of cognitive control within distributed human-AI systems that cannot be explained by automation bias or reliance on algorithms alone. We propose the AIRIS (AI-Augmented Inquiry and Regulation in Hybrid Systems) framework to analyze this dilemma and specify where regulatory intervention can counteract it. AIRIS is a multi-level control allocation architecture specifying the conditions under which epistemic agency can be preserved in hybrid generative systems. Drawing on distributed cognition, cognitive load theory, multimedia learning, and self-regulated learning, it identifies seven interacting mechanisms through which hybrid cognition may become destabilized, from delegation and calibration drift to motivational-affective drift. Five regulatory operators (Anticipate, Interrogate, Reflect, Integrate, and Synthesize) target internal generative engagement at points of emerging instability. The architecture does not itself improve learning; it specifies what must remain in place for genAI-supported work to sustain understanding, whether through instructional design, teacher guidance, or learners'own regulation. We derive testable propositions concerning the seven mechanisms and the five operators, reframing AI augmentation as a problem of control allocation in distributed generative systems. Beyond theory, AIRIS offers a research agenda, a design framework for genAI-integrated learning environments, and a conceptual toolkit for the governance of hybrid human-AI cognition.

Jochen Kuhn, P. Gerjets, Ulrich Trautwein et al. · 0 citations