Skip to content
Open access

Generative AI, Performance, and Learning: A Framework for Comparing Interventions Across Assessment Regimes

Sep 2026 · Computers · Vol 15, pp. 633 · 0 citations · 63 references

TL;DR

A framework for making conditions explicit before intervention effects are compared or synthesized is proposed, which separates what is being assessed, the task, the conditions under which the outcome is produced, and evidence used to verify AI use or non-use.

Abstract

Generative artificial intelligence (AI) can improve students’ work while assistance is available, but better AI-assisted performance does not necessarily show what students have learned or can do independently. The methodological gap is that outcomes collected under different AI-access, timing, and task conditions are often treated as comparable even when they answer different questions. This paper proposes a framework for making those conditions explicit before intervention effects are compared or synthesized. It separates what is being assessed, the task, the conditions under which the outcome is produced, and evidence used to verify AI use or non-use. A retrospective secondary analysis of Wong and Qiu illustrates the problem. For expert-rated originality, unrestricted ChatGPT-4 use outperformed a learner-first strategy on a stuffed-bunny improvement task (+0.78 points), whereas learner-first outperformed unrestricted use on a vocabulary-game task under the study’s no-AI protocol (−0.73; Holm-adjusted p=0.01351). Usefulness showed the same directional pattern, whereas elaboration did not. These results do not establish general creativity, durable learning, or a causal effect of removing AI because task, sequence, and access conditions also changed. For educators and researchers, the practical implication is that AI-assisted performance, independent performance, retention, and transfer are different outcomes and should not be treated as interchangeable without justification. The proposed Assessment-Regime Reporting Profile provides a compact way to document these conditions before findings are generalized, ranked, or pooled.

Read PDF

Similar papers

#human-computer interacti... Preprint Sep 2026

When AI Tutors Speak: Evidence from a Randomized Field Experiment

Students increasingly study alongside generative artificial intelligence (AI), yet unguided access to fluent answers invites cognitive offloading, and there is little evidence on which configurations of AI tutoring produce learning. Two design margins are usually bundled together: pedagogical structure (how the tutor t...

Shi-Hao Yang, Marshall W. van Alstyne, Chrysanthos N. Dellarocas · 0 citations
Review Aug 2026

Evaluation in the Age of AI: Output as Evidence of Learning

This paper argues that the core problem extends beyond academic dishonesty to a deeper misalignment between assessment practices and the learning outcomes they are intended to measure, and highlights the need for alternative assessment models that emphasize process over product.

Md Zarzees Uddin Shah Chowdhury, Samin Khan · 0 citations
Open access Aug 2026

AI-SUPPORTED LEARNING AND HUMAN CAPITAL DEVELOPMENT THROUGH STUDENT-NEGOTIATED INTELLIGENCE

The study introduces the concept of negotiated intelligence, referring to how students regulate AI use while maintaining control over meaning, and highlights the role of contextual intelligence, especially when adapting content to local situations or specific audiences.

Hang Thi Thu Dang, Dao San Tran, Long Ngoc Dinh · 0 citations
Review

Can EdTech Close Learning Gaps? Global Evidence from Digital Interventions *

A systematic review and meta-analysis of randomized controlled trials brings computer-assisted learning platforms and generative AI tools into a common framework under common inclusion criteria and on a common effect-size scale, finding no evidence that generative tools outperform the technologies that preceded them.

Alegria Burneo, Lelys Dinarte-Diaz, C. Lopez et al. · 2 citations
Book Open access Aug 2026

Steering AI Tutors Through System Prompts: A Crossover Study on Self-Regulated Learning and Cognitive Engagement Scaffolds in CS1

The results suggest that the pedagogical behavior of AI tutors may not be easily steered through system prompts alone: embedding established SRL and CE frameworks did not produce detectable improvements on any preregistered outcome in a large, ecologically valid deployment.

Maximilian Georg Barth, Sverrir Thorgeirsson, K. Etemadi et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Beyond"ChatGPT Can Make Mistakes": Designing Interventions to Support Metacognitive Monitoring in AI-Assisted Work

A shared vocabulary, a design space, and evidence that measured monitoring and task performance are separable design targets are contributed, suggesting that measured monitoring and task performance are separable design targets.

M. A. D. Santos, P. Thiesse, Steeven Villa et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.