Metacognition in AI-Supported Second Language Learning: A Systematic Review of Constructs, Measures, and Evidence (2016–2026)
Abstract
Metacognition has informed research on second and foreign language (L2) learning for three decades, while AI tools have supported such learning for nearly as long. Yet, no review has examined how the field conceptualises and measures metacognition under AI mediation. This systematic review identified 60 eligible studies published between 2016 and May 2026. Each study was coded on fourteen dimensions anchored in established taxonomies and appraised with the Mixed Methods Appraisal Tool. Only four eligible studies fall within the pre-ChatGPT portion of the review window, from January 2016 to November 2022; 56 were published during the subsequent three and a half years. Three findings emerge. First, measurement is narrowly concentrated: 98% of the 56 post-ChatGPT studies measured self-reported metacognitive regulation, and 66% relied on self-report alone. Only one study in the full corpus measured metacognitive accuracy. Five post-ChatGPT studies combined an AI-absent comparison with process-level data, and only one of these also measured judgement accuracy. Second, 41% of post-ChatGPT studies used tools that combined two or more functional orientations within a single interface, a configuration not represented by the field’s inherited four-category taxonomy of AI applications. Third, none of the 18 contrast-based studies assigned the regulatory operation to the system in place of the learner. In every design that specified the division of labour, the learner retained that operation; three designs did not specify it. The condition most in need of testing therefore remains untested. Reported findings skew positive, but correlational and qualitative designs account for a disproportionate share of that evidence. The review concludes with a research agenda centred on calibration and process data.