Large Language Model-Based Automated Assessment: A Systematic Review, Taxonomy, and Implications for Personalized Learning
Abstract
Research on large language model (LLM)-based automated assessment (AA) has expanded rapidly. Nevertheless, the literature remains fragmented across contributions, models, implementation configurations, datasets, and evaluation metrics, complicating efforts to identify approaches suitable for personalized learning. This systematic review synthesizes recent research on LLM-based AA and evaluates its implications for delivering personalized learning. Following PRISMA, we conducted a Scopus search for studies published between 2020 and 2025. Two independent reviewers screened the records with strong agreement (Cohen’s $\kappa =0.85$ ), resulting in the inclusion of 65 studies that met the full quality-assessment criteria. The review identifies a sharp increase in LLM-based AA research after 2022. Scoring is the predominant contribution (49 of 65 studies), whereas fully integrated end-to-end pipelines are limited (9 of 65). Decoder-only architectures are prevalent (64 of 65 studies), with GPT-family models used most often. In several studies that conducted within-study comparisons, layered configurations outperformed simpler baselines. However, heterogeneous tasks, datasets, and metrics preclude a cross-study ranking of configurations. Validation maturity varies across the literature: 38.5% of studies are conceptual, while 12.3% include evaluation with real learners. The findings are consolidated into a five-dimensional taxonomy, and a conceptual integration framework is proposed to link LLM-based AA outputs to mastery-oriented adaptive or personalized learning. Validation by three domain experts provided preliminary content-validity evidence for the framework’s axioms (92% agreement) and overall design (94%), resulting in one structural refinement. Overall, LLM-based AA shows promise for scalable assessment and personalized learning, but its effectiveness still needs validation with real learners.