An Efficient Class-Weighted Machine Learning Approach for Early Stroke Detection
Stroke is a significant cause of disability and death across the globe, and there is a need to have trustworthy systems of early prediction to aid in prevention, healthcare, and clinical decision-making. Even though machine learning methods have shown encouraging behaviour in terms of stroke prediction, they can usually be confined to their limits due to the presence of severe imbalance between classes in real-world healthcare data. In this paper, the authors suggest a practical framework of class-weighted machine learning to identify early stroke in a structured clinical dataset. The framework applies systematic exploratory data analysis, strong preprocessing, and model-level control of class imbalance with class-weighted learning, therefore, not relying on synthetic resampling and keeping data intact. Several supervised learning models, such as Logistic Regression, Decision Tree, Random Forest, and Support Vector Machine, are applied and tested within a familiar experimental environment. The accuracy, weighted precision, recall, F1-score, confusion matrix analysis, and ROC-AUC are used to evaluate the performance. It is shown that with experimental results, the class-weighted Random Forest model would perform better, and the accuracy of the model is 94.2% with a weighted F1-score of 95.0% and ROC-AUC of 0.875, which is remarkable because of its discriminative ability even with data imbalance. The suggested solution offers a consistent, interpretable, and clinically viable solution to early stroke prediction risk and can be successfully incorporated into healthcare systems of decision support.