Skip to content
#small language model Review Open access

Conditional validity in LLM-mediated L2 assessment: an argument-based systematic review and meta-analysis

Sep 2026 · Frontiers in Research Metrics and Analytics · Vol 11 · 0 citations · 61 references
Medicine

Abstract

Introduction Large language models (LLMs) are increasingly used for scoring and feedback in second-language (L2) assessment, yet the validity of the resulting interpretations remains contested. This review evaluated when LLM-mediated assessment is psychometrically and educationally defensible using an argument-based validity framework focused on construct representation, reliability and reproducibility, fairness, and washback. Methods Following PRISMA 2020, we systematically searched nine databases and repositories for empirical studies published from January 2022 to December 2025. Fifty-two studies met the inclusion criteria. Quantitative synthesis used REML random-effects models where effects were sufficiently comparable; Pearson correlations, rank correlations, ICC/κ/QWK, and other psychometric indices were otherwise retained on metric-appropriate scales. Washback outcomes were synthesized using Hedges' g. Results Source-level auditing showed that directly comparable Pearson evidence was sparse. Three source-verified L2 studies yielded moderate human-LLM score correspondence (r = 0.66, 95% CI [0.53, 0.75]) with substantial heterogeneity (I2 = 70.8%). Evidence was strongest in rubric-guided, structurally constrained, and human-supervised contexts, whereas construct representation, fairness, operational reproducibility, and generalizability remained underdeveloped. Across 14 studies, LLM-mediated formative feedback produced a small-to-moderate positive effect on revision quality and short-term writing outcomes (g = 0.42, 95% CI [0.27, 0.57]), although heterogeneity was substantial (I2 = 69.8%) and concerns about cognitive offloading, learner agency, and reproducibility persisted. Discussion The findings support a context-dependent interpretation of validity rather than a universal claim about LLM assessment capability. Current evidence supports carefully bounded formative and human-supervised applications but does not justify autonomous high-stakes scoring or broad claims of construct validity, fairness, and generalizability.

Read PDF

Similar papers

Review Open access Sep 2026

Effects of AI-generated feedback on l2 writing: a PRISMA 2020 systematic review of performance, revision, and engagement outcomes

Generative artificial intelligence (GenAI) tools can provide immediate feedback during second-language (L2) writing, but evidence on their effects across writing outcomes remains fragmented. This PRISMA 2020 systematic review synthesized peer-reviewed empirical studies published from 2023 to 2025 that examined GenAI-ge...

Kristian Burhan · 0 citations
#artificial intelligence Review Sep 2026

Beyond AI Literacy: A Structured Review and Exploratory Meta-Analysis of Measures for Competent Generative-AI Use

A structured, seeded review of 24 focal empirical publications, finding no validated individual-level instrument in the focal corpus that tests the full combination of agent scope, permissions, recovery, state isolation, independent review, and evidence-based closure.

D. Véri · 0 citations
Review Open access Oct 2026

Metacognition in AI-Supported Second Language Learning: A Systematic Review of Constructs, Measures, and Evidence (2016–2026)

Metacognition has informed research on second and foreign language (L2) learning for three decades, while AI tools have supported such learning for nearly as long. Yet, no review has examined how the field conceptualises and measures metacognition under AI mediation. This systematic review identified 60 eligible studie...

Hong Yi, Qiang Chen, Zhuo Wang · 0 citations
Open access Aug 2026

From policy debate to empirical evidence: a proof-of-concept bibliometric and LLM-assisted abstract-level content analysis of OTC hearing aid research before and after FDA regulation.

OBJECTIVE To characterise temporal patterns in over-the-counter (OTC) hearing-aid research before and after U.S. Food and Drug Administration (FDA) implementation of the OTC category (effective October 2022), using bibliometrics and large language model (LLM)-assisted abstract-level annotation. DESIGN Bibliometric an...

Chang-Geng Mo, V. Manchaiah, J. Wasmann et al. · 0 citations
Open access Sep 2026

AI Psychometrics: Evaluating Measurement Invariance in Science Self-Efficacy Using Large Language Models

The present study (1) examined sex- and language-based measurement invariance of the science self-efficacy scale using PISA 2006 and 2015 U.S. data (N = 11,323) and (2) evaluated six large language models (ChatGPT, DeepSeek, Grok, Copilot, MetaAI, Gemini) in generating invariance approximations from prompt content and...

Onur Ramazan, D. Lee · 0 citations
Aug 2026

Assessing AI-Mediated Interviewing Quality: A Theory-Grounded Comparison of Models for Evaluation Data Collection

As large language models are increasingly used in evaluative settings for their conversational potential in interviews, systematic frameworks for assessing their methodological quality are still lacking. Drawing on evaluation, qualitative interviewing, and validity theories, this study develops a theory-grounded framew...

Ali Safarnejad, Hippolyte Lefebvre · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.