Skip to content
Conference Open access

Machine Learning for Stroke Onset Risk Prediction and Model Comparison

2026 · ITM Web of Conferences · 0 citations · 2 references

Abstract

Stroke is a leading cause of death and disability worldwide, and early risk prediction is essential for reducing its public health burden. In this paper, six machine learning models were developed and thoroughly compared. An analysis of stroke risk using the publicly available Kaggle Stroke data set. The article discusses the Prediction Dataset and the use of Logistic Regression, Random Forest, and Support Vector methods. The performance of Machine, K-Nearest Neighbors, Decision Tree, and XGBoost was assessed by means of the Accuracy, Precision, Recall, F1-score, and AUC-ROC metrics. Because class weighting was used to deal with the severe data imbalance, it was found that. Since Logistic Regression, Random Forest, and XGBoost all gave AUC-ROC values above 0.80, they therefore meet the clinically acceptable threshold for effectiveness. Since regression gave the highest Recall of 0.84, it is clear that regression is superior. The performance of identifying high-risk patients is discussed in connection with feature importance. From the analysis it was clearly established that age, average glucose level, and body mass index were the most important predictors, hence the study properly validates this. The paper discusses the feasibility of using machine learning for stroke risk prediction and therefore gives a very useful reference for clinical auxiliary screening.

Read PDF