Heterogeneous stacking framework with adaptive Borderline-SMOTE for imbalanced binary classification: a multi-institutional validation study
Abstract
Binary classification in imbalanced tabular datasets remains a significant challenge in machine learning, as conventional risk-stratification models exhibit limited discriminative performance and fail to capture nonlinear interactions among heterogeneous features. Existing approaches often suffer from three critical limitations: model-selection uncertainty across heterogeneous data distributions, systematic bias toward majority classes when training on imbalanced datasets, and insufficient interpretability for high-stakes decision-making applications. This work proposes a heterogeneous stacking ensemble framework that integrates five complementary base learners—Random Forest, XGBoost, LightGBM, CatBoost, and a Multi-Layer Perceptron—through an L2-regularized logistic regression meta-learner trained on out-of-fold predictions. To address class imbalance, we incorporate an adaptive Borderline-SMOTE oversampling strategy within a leakage-free cross-validation pipeline that concentrates synthetic sample generation in high-difficulty borderline regions. The framework is developed on a multi-institutional dataset of 4,127 instances with 28 features and externally validated on two independent cohorts (n=612, n=489). The proposed approach achieves an AUC of 0.892 (95% CI: 0.876–0.908) on the internal test set, outperforming the strongest single learner by 2.8% and four state-of-the-art ensemble baselines by margins of 1.3% to 3.4% in AUC. External validation demonstrates robust transportability with AUCs of 0.871 and 0.858 respectively. Decision-curve analysis confirms superior net benefit across threshold ranges of 10–40%. Comprehensive ablation studies covering feature-group contributions, base-learner combinations, resampling strategies, and meta-learner architectures provide empirical guidance for ensemble design in imbalanced classification tasks. Model decisions are explained at global and instance levels using SHAP value analysis, revealing the dominant predictive features and demonstrating consistency with domain knowledge. The proposed framework offers an accurate, interpretable, and generalizable solution for high-stakes tabular classification problems with inherent class imbalance.