LLM-Generated Feedback in L2 Writing: A Scoping Review
Abstract
The release of ChatGPT in November 2022 transformed second language (L2) writing instruction and led to rapid growth in research on large language model (LLM)-generated feedback; however, no synthesis has mapped this literature in terms of feedback quality, learner uptake, and pedagogical integration. This scoping review examines 185 empirical studies published between November 2022 and March 2026 that were identified through a Scopus search (n = 283 screened) and analysed using a systematic keyword-based charting framework applied to full abstracts, with full-text analysis of 35 studies. The review identifies four major patterns: (1) comparative AI–human feedback research dominates the literature (27.6%); (2) content-level feedback remains underexplored (13.5% of studies); (3) learner uptake is rarely measured as a primary outcome; and (4) LLM feedback is broadly comparable to teacher feedback for surface-level errors but weaker for content and argumentation, while learner perceptions often exceed demonstrated performance outcomes. Hybrid AI–teacher models show promising but underexamined potential, accounting for only 12.4% of the literature. The field shows a focus on perceptions rather than learning outcomes, an apparent tendency toward positive-results reporting, and no clear teaching models. This study proposes a typology of LLM feedback functions and outlines a research agenda focused on uptake, longitudinal outcomes, and hybrid AI–teacher integration.