Skip to content
Open access

Academic Dropout Prediction Using Large-Scale Static Institutional Data: A Multi-Scenario Machine Learning Study

Sep 2026 · Technologies · 0 citations · 23 references

Abstract

Academic dropout is costly for students and institutions, and support actions are most effective early, when little information beyond enrollment records is available. However, the most informative predictors are derived from academic trajectory data, such as grades and course progression, which only become available after students have completed one or more terms. This study investigates dropout prediction using supervised machine learning (ML) applied to records of 66,820 students across 45 undergraduate programs of the State University of Maringá, Brazil (2002–2022), trained exclusively on static enrollment-time variables, since these were the only institutional data available. Although this restricts the information the models can use, it also allows students who may be at risk to be flagged very early. Decision Tree, Random Forest, and eXtreme Gradient Boosting (XGBoost) models were evaluated in ten training configurations, covering the complete dataset, the complete-generations subset, a temporal cohort split, academic centers, individual programs, and resampled variants. Evaluation was based on per-class precision, recall, and F1-score, together with threshold-free and calibration measures (PR-AUC, ROC-AUC, and Brier score). XGBoost performed best in the institution-wide configurations, reaching a dropout recall of 0.64 at a precision of 0.51 with undersampling and, without resampling, a PR-AUC of 0.586 against a dropout prevalence of 0.364, which corresponds to 1.61 times the performance of random ranking. Under the temporal split, recall at the default threshold fell from 0.37 to 0.18, while ROC-AUC changed little (0.708 to 0.679); this loss reflects miscalibration under changing dropout prevalence and can be corrected without retraining the model. The main contributions are an evaluation methodology for enrollment-time dropout prediction that controls information leakage and includes temporal validation, and a deployment protocol that uses the model outputs to prioritize student support when follow-up capacity is limited.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.