Hybrid Machine Learning Models for Predictive Analytics and Intelligent Decision-Making in Complex Data Environments
Abstract
Hybrid machine learning offers a flexible approach for predictive analytics and intelligent decision-making in heterogeneous data environments. This study developed and evaluated Logistic Regression, Random Forest, soft voting, and weighted voting using a publicly available educational dataset comprising 4,424 student records, 36 predictors, and three outcome classes: dropout, enrolled, and graduate. Data preprocessing included categorical encoding, numerical scaling, class-weighted learning, and stratified training-testing separation. The models were compared using accuracy, balanced accuracy, macro precision, macro recall, macro F1-score, receiver operating characteristic area under the curve, log loss, and calibration measures. The weighted-voting ensemble achieved the best overall performance, with 76.72% accuracy, 71.89% balanced accuracy, 71.81% macro F1-score, and an area under the curve of 0.904. Performance improved from a macro F1-score of 55.54% with baseline predictors to 65.91% after first-semester information and 71.81% after second-semester information were added. Academic progression indicators, tuition-fee status, age at enrolment, debtor status, course, and gross domestic product were influential predictors. Probability thresholds supported high-priority intervention, moderate-risk monitoring, and uncertain-case review. The enrolled category remained the most difficult outcome to classify because of imbalance and transitional characteristics. Hybrid ensembles provided accurate, interpretable, and decision-oriented predictions, although external validation and prospective evaluation are required before institutional deployment.