Sep 2026· Frontiers in Research Metrics and Analytics· Vol 11· 0 citations· 61 references
Medicine
Abstract
Introduction Large language models (LLMs) are increasingly used for scoring and feedback in second-language (L2) assessment, yet the validity of the resulting interpretations remains contested. This review evaluated when LLM-mediated assessment is psychometrically and educationally defensible using an argument-based validity framework focused on construct representation, reliability and reproducibility, fairness, and washback. Methods Following PRISMA 2020, we systematically searched nine databases and repositories for empirical studies published from January 2022 to December 2025. Fifty-two studies met the inclusion criteria. Quantitative synthesis used REML random-effects models where effects were sufficiently comparable; Pearson correlations, rank correlations, ICC/κ/QWK, and other psychometric indices were otherwise retained on metric-appropriate scales. Washback outcomes were synthesized using Hedges' g. Results Source-level auditing showed that directly comparable Pearson evidence was sparse. Three source-verified L2 studies yielded moderate human-LLM score correspondence (r = 0.66, 95% CI [0.53, 0.75]) with substantial heterogeneity (I2 = 70.8%). Evidence was strongest in rubric-guided, structurally constrained, and human-supervised contexts, whereas construct representation, fairness, operational reproducibility, and generalizability remained underdeveloped. Across 14 studies, LLM-mediated formative feedback produced a small-to-moderate positive effect on revision quality and short-term writing outcomes (g = 0.42, 95% CI [0.27, 0.57]), although heterogeneity was substantial (I2 = 69.8%) and concerns about cognitive offloading, learner agency, and reproducibility persisted. Discussion The findings support a context-dependent interpretation of validity rather than a universal claim about LLM assessment capability. Current evidence supports carefully bounded formative and human-supervised applications but does not justify autonomous high-stakes scoring or broad claims of construct validity, fairness, and generalizability.
Generative artificial intelligence (GenAI) tools can provide immediate feedback during second-language (L2) writing, but evidence on their effects across writing outcomes remains fragmented. This PRISMA 2020 systematic review synthesized peer-reviewed empirical studies published from 2023 to 2025 that examined GenAI-ge...
Kristian Burhan· JRTI (Jurnal Riset Tindakan...· 0 citations
A structured, seeded review of 24 focal empirical publications, finding no validated individual-level instrument in the focal corpus that tests the full combination of agent scope, permissions, recovery, state isolation, independent review, and evidence-based closure.
Metacognition has informed research on second and foreign language (L2) learning for three decades, while AI tools have supported such learning for nearly as long. Yet, no review has examined how the field conceptualises and measures metacognition under AI mediation. This systematic review identified 60 eligible studie...
Hong Yi, Qiang Chen, Zhuo Wang· Journal of Intelligence· 0 citations
OBJECTIVE
To characterise temporal patterns in over-the-counter (OTC) hearing-aid research before and after U.S. Food and Drug Administration (FDA) implementation of the OTC category (effective October 2022), using bibliometrics and large language model (LLM)-assisted abstract-level annotation.
DESIGN
Bibliometric an...
Chang-Geng Mo, V. Manchaiah, J. Wasmann et al.· International Journal of Aud...· 0 citations
The present study (1) examined sex- and language-based measurement invariance of the science self-efficacy scale using PISA 2006 and 2015 U.S. data (N = 11,323) and (2) evaluated six large language models (ChatGPT, DeepSeek, Grok, Copilot, MetaAI, Gemini) in generating invariance approximations from prompt content and...
Onur Ramazan, D. Lee· Psychology International· 0 citations
As large language models are increasingly used in evaluative settings for their conversational potential in interviews, systematic frameworks for assessing their methodological quality are still lacking. Drawing on evaluation, qualitative interviewing, and validity theories, this study develops a theory-grounded framew...
Ali Safarnejad, Hippolyte Lefebvre· American Journal of Evaluati...· 0 citations