Skip to content
Preprint

Experimental Evidence on the Learning Impact of Generative AI

Jul 2026 · 1 citation
Economics Computer Science

TL;DR

Evidence is found for two mechanisms behind the learning gains: students shift time away from drafting text and toward reading and searching for information, and they report greater learning enjoyment.

Abstract

We study how generative AI affects student learning in a randomized experiment. In proctored, in-person sessions, undergraduates learn about an unfamiliar topic and write an analytical essay with or without access to off-the-shelf generative AI, then complete unaided assessments immediately and one week later. We measure learning with knowledge tests (factual and conceptual understanding) and open-ended essays (higher-order skills). AI access raises immediate test scores by 0.27 standard deviations. These gains persist one week later. Essay quality, by contrast, changes little while students have AI access but improves in style and relevance one week later, when students write unaided. These delayed gains are larger among augmentation users-who use AI to explain concepts rather than generate text-whereas automation users'short-run quality gains vanish once AI is removed. We find evidence for two mechanisms behind the learning gains: students shift time away from drafting text and toward reading and searching for information, and they report greater learning enjoyment.

View source

Similar papers

Preprint Jul 2026

Student Evaluation of Repeated AI Feedback Across a Semester of Writing

Generative AI is increasingly used for feedback in higher education, but evidence from repeated classroom use remains limited. This short paper analyses 2988 reflective essay-feedback-appraisal instances from 283 Estonian bachelor students across one semester. Students obtained and assessed feedback from a self-selected AI tool using a uniform prompt. The present analysis of the anonymized text corpus covers essay content, AI feedback, and its perceived helpfulness. Students found feedback helpful and actionable more often than not; about a tenth thought AI unhelpful, more so towards the end of the semester. We also analyzed essay reflection depth, and used a validated AI text classifier to estimate the share of essays that could be treated as likely unaided student writing. The study contributes descriptive classroom evidence on integration of AI feedback - a fast and scalable way to provide immediate writing advice, but not a self-contained route to better reflection. Benefits depend on whether students learn to use AI selectively and critically, without sliding into over-use harmful for the learning process.

Andres Karjus, Janika Leoste, Tiia Õun · 0 citations
Preprint Jul 2026

Generative AI Availability, Grades, and Student Satisfaction at a Large University

It is found that there is no significant differential effect of GenAI availability on grades overall or among previously lower-performing students, and the findings temper concerns that GenAI inflates grades and reduces students's satisfaction.

J. Dumlao, Meng Wang, Zhonghan Xie et al. · 0 citations
#generative ai Review Open access Aug 2026

Impact of Generative Artificial Intelligence Use on the Academic Performance of Undergraduate Students in Canada

Generative artificial intelligence (AI) has entered everyday undergraduate study faster than the evidence base has kept up with it, and the evidence that does exist points in opposite directions. This paper synthesizes peer-reviewed and carefully delimited contextual research on how generative AI use relates to undergraduate academic performance, with attention to what Canadian universities can already act on. Searches of Google Scholar, Scopus, Web of Science, ERIC, and ScienceDirect covering 2022 to 2026 produced a two-tier evidence base: Tier 1 peer-reviewed empirical and synthetic studies of student learning outcomes, and Tier 2 supplementary sources admitted under stated justifications, including contextual Canadian and international surveys, one secondary-school field experiment, and one preprint mechanistic study. Experimental syntheses report medium to large short-term gains when ChatGPT is built into instruction. Survey work points the other way: frequent unstructured use tracks with procrastination, self-reported memory problems, and slightly lower grades, and unrestricted access during practice has been shown to depress later unaided performance. Purpose of use reconciles most of that disagreement, because a tool that scaffolds thinking behaves very differently from one that replaces it. Canadian peer-reviewed studies document heavy campus adoption and considerable student ambivalence about integrity and learning, yet almost none link purpose-differentiated use to measured performance. Closing that gap matters for assessment redesign, AI-literacy programming, and the credibility of the credentials Canadian universities issue.

Pragalvha Sharma · 0 citations
Aug 2026

AI Literacy as Experimental Practice: Students as Investigators

A three-week midterm project embedded in an undergraduate “AI-for-all” course investigated whether AI literacy can be taught to undergrads, and shows any user how to test an AI system rather than trust it blindly.

Amarda Shehu, Adonyas Ababu, Asma Akbary et al. · 0 citations
Open access Jul 2026

A Metacognitive Blind Spot: Student Comprehension, AI Reliance, and the Conceptual Difficulty Gap

This exploratory study investigates the relationship between student metacognition, use of artificial intelligence, and empirical performance within a multidisciplinary course on AI. Using data from a sample of college-age, full-time undergraduate students (averaging 18 participants per assessment) enrolled in an in-person junior seminar at a Midwestern U.S. university, we correlate student self-assessments with standard readability metrics (e.g., Flesch–Kincaid), L2SCA metrics, and objective assessment outcomes, analyzing how learners evaluate their own comprehension and how they deploy AI tools in response to the complexity of 21 reading assignments over 8 weeks. We find that students’ perceptions of linguistic difficulty correlate with classical readability scores, but their perceptions do not predict their success nor does their engagement with assistive AI. The results suggest that students utilize generative AI tools as a habitual baseline rather than a strategic response to difficult material. We argue that while students can identify surface-level linguistic friction, they fail to recognize deep conceptual hurdles, leading to a false sense of mastery that neither their intuition nor their AI assistants appear to mitigate. We propose that a quantifiable metric, the Conceptual Difficulty Gap (CDG), may be useful for identifying a class of texts that syntactically appear to be simple, but consistently trigger performance failures. Crucially, we uncover a possible metacognitive blind spot: student self-ratings of difficulty are negatively correlated with this gap, implying that student assessments of difficulty are not based on actual conceptual difficulty. Furthermore, self-reported AI reliance shows no correlation with the gap, indicating that students may not be strategically deploying generative AI tools to mitigate conceptual difficulty.

Igor Crk, E. Gultepe · 0 citations