Skip to content
Open access

Comparative Evaluation of Machine Learning Models for Early Prediction of Student Academic Performance Using the Open University Learning Analytics Dataset

Sep 2026 · International Journal of Interactive Mobile Technologies (ijim) · 0 citations

TL;DR

This study develops and comparatively evaluates four supervised machine learning models for multi-class prediction of final student outcome on the Open University Learning Analytics Dataset (OULAD), demonstrating that aggregate accuracy substantially overstates the usefulness of these models for risk detection and that the fail class remains the principal bottleneck.

Abstract

Early identification of students at risk of academic failure is a central problem in learning analytics, yet the multi-table, behaviorally heterogeneous, and class-imbalanced nature of educational data complicates reliable prediction. This study develops and comparatively evaluates four supervised machine learning models—logistic regression, support vector machine (SVM) with a radial basis function kernel, random forest, and extreme gradient boosting (XGBoost)—for multi-class prediction of final student outcome on the Open University Learning Analytics Dataset (OULAD), comprising 32,593 anonymized student records and more than ten million virtual-learning-environment interaction logs. A reproducible, leakage-controlled feature-engineering pipeline is proposed: final examination records are excluded, and virtual-learning-environment activity is restricted to the first 120 days of each presentation, yielding 85 demographics, assessment, and behavioral features. All models are trained on an identical stratified 80/20 split and assessed within a unified evaluation framework using accuracy, weighted and macro F1-score, balanced accuracy, per-class recall, and stratified cross-validation. XGBoost achieved the strongest overall performance (accuracy 0.7302, macro F1 0.6658, balanced accuracy 0.6555, and cross-validated accuracy 0.7314), followed closely by random forest and logistic regression, while SVM, despite the weakest aggregate scores, produced the highest recall for the practically critical Fail class (0.4026). The analysis demonstrates that aggregate accuracy substantially overstates the usefulness of these models for risk detection and that the fail class remains the principal bottleneck. A deployed Streamlit prototype illustrates operationalization. The contribution lies in the leakage-aware early-window pipeline and the imbalance-sensitive, algorithm-aware comparison that together expose where predictive value is gained and lost.

Read PDF

Similar papers

Open access Aug 2026

Early identification of at-risk students in online learning environments: A learning analytics approach using machine learning models

Online learning environments generate temporally ordered traces that can support early academic-risk screening, but credible claims require an explicit prediction point and strict control of post-outcome information. This study evaluates whether information available by day 28 of an Open University module presentation...

Verda Gizem Oğul, Y. S. Balcıoğlu · 0 citations
Open access Sep 2026

Student Performance Prediction Using Machine Learning: A Comparative Analysis of Decision Tree, Random Forest, and Optimized XGBoost

Student performance prediction has become an important application of Machine Learning in educational data mining, enabling institutions to identify academically weak students at an early stage and provide timely academic support. Accurate prediction of student performance helps educators implement personalized learnin...

K. S. Sangeetha, Tulasi Miryala · 0 citations
Open access Aug 2026

Interpretable Machine Learning for Early Detection of Academically At-Risk Students Using Behavioral, Socio-Demographic and Learning-Related Factors

The increasing adoption of learning analytics in higher education has encouraged the development of machine learning models for the early detection of academically at-risk students. However, many predictive models emphasize accuracy while providing limited interpretability for educators and academic advisors. This stud...

Vadlya Maarif · 0 citations
Review Open access Sep 2026

An explainable machine learning framework for early student dropout risk prediction and stratification

Student dropout remains a persistent challenge in higher education institutions, affecting academic continuity, institutional performance, and long-term socioeconomic outcomes. Early identification of at-risk students enables timely intervention and improved retention strategies. This study proposes an explainable mach...

Vinayak Hegde, Anuvarshini Palaniswamy, Doyel Bhar et al. · 0 citations
Open access Sep 2026

Academic Dropout Prediction Using Large-Scale Static Institutional Data: A Multi-Scenario Machine Learning Study

Academic dropout is costly for students and institutions, and support actions are most effective early, when little information beyond enrollment records is available. However, the most informative predictors are derived from academic trajectory data, such as grades and course progression, which only become available a...

Rômulo Barreto Mincache, L. Catharin, Lucas de Oliveira Teixeira et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.