Skip to content
Review Open access

Large Language Model-Based Automated Assessment: A Systematic Review, Taxonomy, and Implications for Personalized Learning

2026 · IEEE Access · Vol 14, pp. 149761-149791 · 0 citations · 112 references

Abstract

Research on large language model (LLM)-based automated assessment (AA) has expanded rapidly. Nevertheless, the literature remains fragmented across contributions, models, implementation configurations, datasets, and evaluation metrics, complicating efforts to identify approaches suitable for personalized learning. This systematic review synthesizes recent research on LLM-based AA and evaluates its implications for delivering personalized learning. Following PRISMA, we conducted a Scopus search for studies published between 2020 and 2025. Two independent reviewers screened the records with strong agreement (Cohen’s $\kappa =0.85$ ), resulting in the inclusion of 65 studies that met the full quality-assessment criteria. The review identifies a sharp increase in LLM-based AA research after 2022. Scoring is the predominant contribution (49 of 65 studies), whereas fully integrated end-to-end pipelines are limited (9 of 65). Decoder-only architectures are prevalent (64 of 65 studies), with GPT-family models used most often. In several studies that conducted within-study comparisons, layered configurations outperformed simpler baselines. However, heterogeneous tasks, datasets, and metrics preclude a cross-study ranking of configurations. Validation maturity varies across the literature: 38.5% of studies are conceptual, while 12.3% include evaluation with real learners. The findings are consolidated into a five-dimensional taxonomy, and a conceptual integration framework is proposed to link LLM-based AA outputs to mastery-oriented adaptive or personalized learning. Validation by three domain experts provided preliminary content-validity evidence for the framework’s axioms (92% agreement) and overall design (94%), resulting in one structural refinement. Overall, LLM-based AA shows promise for scalable assessment and personalized learning, but its effectiveness still needs validation with real learners.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.