Using counterfactual explanations to enable prescriptive analytics in educational data mining projects
Abstract
As machine learning (ML) models are increasingly used to support decision making in higher education, there is a growing need to move beyond accurate prediction toward explanations that enable meaningful intervention. This paper examines counterfactual explanations (CFEs) as prescriptive complements to risk prediction in educational data analytics. While predictive models can flag students at risk, their practical value is limited without actionable guidance. CFEs address this need by proposing small, feasible changes that may shift a prediction from at-risk to not at-risk. Three approaches are compared on the Student Insomnia and Educational Outcomes (SIEO) dataset using a LightGBM classifier as the predictive backbone: DiCE (genetic search), the method of Wachter et al. (optimisation with proximity penalties), and a FACE-lite approximation that follows manifold-constrained paths via nearest-neighbour graphs. Demographic variables (eg gender, year of study) are treated as immutable; behavioural and psychological features (eg sleep duration, fatigue, stress) are allowed to vary. Counterfactual quality is assessed by validity, proximity, sparsity, and plausibility. Results indicate that improving sleep, reducing fatigue, and lowering stress frequently appear as pathways for altering model outputs. The comparison suggests distinct trade-offs: DiCE tends to increase diversity with occasional plausibility concerns; Wachter emphasises minimal changes with more repetitive adjustments; FACE-lite balances plausibility with structural constraints. Taken together, these findings point to CFEs as a practical complement to predictive analytics, offering individualised recommendations that can inform advising and student support. This article is also included in The Business & Management Collection which can be accessed at https://hstalks.com/business/.