Skip to content

An Explainable Machine Learning Framework for Early Mental Health Risk Detection among University Students in Fragile and Conflict-Affected Settings

· 0 citations · 2 references

TL;DR

An explainable machine-learning framework for early mental health risk stratification among university students in FCAS contexts is developed that combines demographic, academic, socioeconomic, psychological, and contextual indicators within a structured modelling pipeline designed to support transparency, calibration, and practical decision-making.

View source

Similar papers

Open access Aug 2026

A Transparent AI-Driven Ensemble Learning Approach for Predicting Mental Health Risks among University Students in Kenya

This study addresses the escalating mental health crisis among university students in Kenya, where psychological distress, driven by academic, financial, social, and transitional pressures, is increasingly prevalent yet under-prioritized in public health discourse. While global studies report distress rates exceeding 75% and Kenyan prevalence exceeds 40%, there remains a critical lack of localized, data-driven models to enable early detection in sub-Saharan African university settings. To address this gap, we developed transparent ensemble machine learning models, Random Forest (RF) and Extreme Gradient Boosting (XGBoost), to predict mental health distress levels (low, moderate, high) among 1,128 students across universities in Tharaka Nithi County, Kenya. Using a structured questionnaire incorporating demographic, academic, financial, psychosocial, and Quarter-Life Crisis (QLC) factors, we trained the models on 70% of the data and evaluated them on the remaining 30%. Both models achieved high performance: RF and XGBoost attained 98.8% accuracy (95% CI: 96.9–99.7%), Cohen’s Kappa of 0.978, and multi-class AUCs of 1.000 (RF) and ≥0.997 (XGBoost). Class-specific metrics revealed near-perfect precision, recall, and F1-scores; for high distress, both models achieved 100% sensitivity and specificity, with F1-scores of 1.000 (RF) and 0.988 (XGBoost). SHAP and feature importance analyses identified “Quarter-Life Crisis,” faculty of study (especially Science, Technology, Nursing, and Engineering), personal/mental health history, financial stress, and residence as top predictors. These findings support the use of interpretable AI for equitable, early-risk screening. However, careful attention must be paid to ethical considerations, including data privacy, potential algorithmic bias, and mental health stigma. We recommend integrating such models into university wellness platforms to enable proactive, targeted support for at-risk students.

Victor Lumumba, Dennis Muriithi, Monicah Oundo · 0 citations
Open access Jul 2026

Explainable XGBoost Early-Warning Framework for Academic Stress-Based Student Mental Health Risk Mapping

Existing university mental health monitoring often depends on voluntary help-seeking or manual questionnaire interpretation, which may delay early support for students experiencing academic stress. This study proposes an explainable XGBoost-based early-warning framework for non-clinical mapping of student mental health risk from academic stress indicators. The single-site dataset comprised 1,002 anonymized student records from Universitas Muria Kudus. K-Means clustering was used to transform DASS-21 depression, anxiety, and stress scores into low, moderate-, and high-risk categories, while XGBoost predicted the cluster-derived labels using seven single-item academic stress indicators and engineered aggregate and interaction features. On a stratified hold-out testing set of 201 records, the model achieved weighted precision, recall, and F1-score values of 0.8907, 0.8905, and 0.8906, respectively, with class-level F1-scores of 0.9109 for low risk, 0.8900 for moderate risk, and 0.8713 for high risk. Additional ablation, clustering sensitivity, subgroup, threshold, and SHAP stability analyses were conducted to strengthen robustness and interpretability. The findings show that cumulative academic stress and interaction features involving parental expectations, exam anxiety, and learning-method adaptation were consistently influential predictors. The framework is intended to support early institutional prioritization and counseling referral, not clinical diagnosis. Generalization remains limited by the single-institution sample and the use of single-item academic stress indicators; therefore, local retraining and recalibration are required before institutional deployment, including implementation of the Streamlit prototype.

Supriyono, H. Firmansyah, Soni Adiyono · 0 citations
Jul 2026

AI-Powered Student Mental Health Analytics and Academic Prediction System: A Supervised Machine Learning and Explainable AI Approach

Student mental health difficulties such as stress, anxiety, depression and poor sleep frequently go unnoticed until they have already affected attendance, grades and personal wellbeing, because traditional counselling depends on a student voluntarily seeking help. This paper presents an AI-Powered Student Mental Health Analytics and Early Intervention System, a role-based web platform that combines validated psychological screening instruments — the PHQ-9, GAD-7 and Perceived Stress Scale (PSS) — with academic and lifestyle indicators such as CGPA, attendance, sleep duration, exercise and social support to build a comprehensive student risk profile. Collected data is cleaned and engineered before being passed through an ensemble of supervised learning models — Random Forest, Gradient Boosting, XGBoost and a Voting Classifier — which jointly classify each student's mental health status as Low, Moderate or High Risk and estimate likely academic performance. SHAP (SHapley Additive exPlanations) is applied so that every prediction is interpretable to the counsellors who act on it. The system is implemented with Django and Python on the backend, MySQL for persistence, and HTML/CSS/Bootstrap/JavaScript with Chart.js on the frontend, and provides three dedicated modules — Student, Counselor and Administrator — each with role-specific dashboards, along with dedicated reporting, analytics, notification and security modules. High-risk classifications trigger automatic counsellor alerts, while an institution-wide analytics dashboard surfaces department- and cohort-level trends. Functional testing across authentication, assessment, prediction, reporting, analytics, counselling and notification confirmed correct end-to-end operation, and the ensemble classifier achieved strong performance on accuracy, precision, recall and F1-score. The system demonstrates that combining validated clinical screening tools with explainable ensemble learning can support proactive, scalable and transparent early intervention in educational institutions.

Sagara C P, Mohammed Zaid, T. Vasudev · 0 citations
Open access Jul 2026

Identification and validation of an explainable screening model of college students' mental health with therapeutic strategy implications.

Mental health issues, especially depressive symptoms, among young adults represent a public health challenge. Conventional psychological assessment tools have limited sensitivity and specificity for identifying individuals at risk. This study aims to develop an explainable machine learning-based model to stratify concurrent depression risk in young adults. This study included 100,257 college students and collected mental health variables including depression, anxiety, resilience, parent-child relationship, and duration of mobile phone usage. The screening capabilities of 13 machine learning algorithms were systematically evaluated and compared. The SHapley Additive exPlanations (SHAP) framework was employed for the interpretability of the final model. The median scores for parent-child relationship, resilience, anxiety, and mobile phone usage time was 42.0, 28.0, 1.0 and 28.0, respectively. Among the 13 machine learning algorithms, the XGBoost model demonstrated superior performance. The final multivariate screening model achieved an area under the curve (AUC) of 0.887, a sensitivity of 0.787, a specificity of 0.830, and an accuracy of 0.816 in classifying young adults' concurrent depression risk. The SHAP analysis showed the importance of each variable: anxiety (2.303) > resilience (0.774) > parent-child relationship (0.708) > mobile phone usage time (0.411). The final multivariate model exhibited stable performance during cross-validation (AUC = 0.885 ± 0.032), significantly better than the single-variable model (P < 0.001) and better screening reliability (Brier score 0.153). The final multivariate XGBoost model provides a highly accurate and interpretable approach for young adults' depression risk stratification. As the model was developed using cross-sectional data collected during the COVID-19 campus lockdown, prospective validation is required before clinical deployment. Notably, anxiety level emerged as the most influential risk factor, and resilience demonstrated a significant protective effect.

Ke-Xuan Liu, Zhe Li, Yu-Yu Zhao et al. · 0 citations
Open access Aug 2026

An explainable hybrid of classical and quantum support vector machine models for maternal mental health risk prediction using multicountry data

Maternal mental health disorders remain a major public health problem, particularly in low- and middle-income countries, where limited mental health resources and inadequate screening systems lead to late diagnosis and intervention. Recent advances in artificial intelligence have demonstrated great potential in improving the prediction of early maternal mental health risks. However, most current methods are limited to traditional machine learning models and lack interpretability. This study introduces the explainable Hybrid Classical-Quantum Support Vector Machine (SVM) framework for the prediction of maternal mental health risk using multicountry psychosocial and clinical data from Uganda and Pakistan. The proposed framework integrates classical SVM learning with quantum-enhanced feature representations, while SHapley Additive exPlanations (SHAP) are leveraged to increase the model transparency and clinical interpretability. A comparative evaluation was performed against Random Forest, XGBoost, LightGBM, CatBoost, Stacking Ensemble, Classical SVM, and standalone Quantum SVM using Accuracy, Precision, Recall, F1-score, Receiver Operating Characteristic Area Under the Curve (ROC-AUC) and Precision-Recall Area Under the Curve (PR-AUC). We performed two experimental settings: one using the full set of features and another using a set of features without leakage, i.e. excluding variables that directly contribute to the EPDS-derived outcome. Under the complete feature set, the proposed Hybrid Classical-Quantum SVM achieved 99.86% accuracy, 99.80% F1-score, 99.92% ROC-AUC and 99.91% PR-AUC, while maintaining competitive performance after the leakage-free evaluation. The SHAP analysis provided enhanced model interpretability, identifying clinically meaningful demographic, obstetric, socioeconomic and psychosocial factors associated with maternal mental health risk. Our results demonstrate that hybrid classical-quantum learning offers a competitive and explainable paradigm for predicting maternal mental health risk and highlights the complementary nature of quantum feature representations in healthcare analytics.

Shallon Ahimbisibwe, Emmanuel Ahishakiye, S. Maling et al. · 0 citations
Open access Jul 2026

An Evidence-Based Integrated Framework for College Student Mental Health Monitoring and Precision Intervention Using Multi-Source Big Data

College student mental health issues are rising, but traditional assessments lack timeliness and objectivity. Leveraging campus digitalization, this study proposed an end-to-end “prediction–interpretation–intervention” framework addressing multisource data heterogeneity, poor interpretability, and ethical concerns. A macro–micro dual-scale pipeline integrated psychological, behavioral, and physiological data. An evidence-based model combining causal inference and attention mechanisms boosted prediction accuracy and explainability; t-distributed stochastic neighbor embedding and clustering revealed three risk subgroups, enabling tiered interventions. The model achieved 0.812 average precision, remained robust under data perturbations, and pilot interventions significantly reduced depressive symptoms and improved health behaviors in high-risk students. This framework bridged data-driven prediction and actionable support, offering a scalable, ethical solution for campus mental health management.

Lanwen Wang, Shiming Shen, Feier Chen · 0 citations