Aug 2026· GSC Advanced Research and Reviews· 0 citations
TL;DR
Male sex was a statistically significant independent predictor of heart disease after controlling for other clinical variables and the findings support sex-specific screening and preventive strategies for high-cholesterol male patients and demonstrate the value of interpretable machine learning models for clinical decision support.
Abstract
Cardiovascular disease remains a major cause of mortality and economic burden in the United States. This study developed and evaluated supervised machine learning models to predict heart disease risk from routine clinical variables and tested whether male patients with high cholesterol have higher odds of heart disease than female patients. Using a publicly available clinical dataset of 918 patients, the Cross-Industry Standard Process for Data Mining (CRISP-DM) framework guided exploratory analysis, data cleaning, median imputation of irregular cholesterol values, and feature selection using principal component analysis (PCA) and SelectKBest. Logistic Regression, Support Vector Machine (SVM), and a soft voting ensemble were trained and tuned using grid search with stratified 5-fold cross-validation. The ensemble achieved the highest predictive performance (accuracy = 0.940, F1 = 0.950, receiver operating characteristic area under the curve [ROC-AUC] = 0.958). Logistic Regression achieved comparable performance (accuracy = 0.929, F1 = 0.940, ROC-AUC = 0.958) and was selected for hypothesis testing because of its interpretability. Among patients with cholesterol ≥240 mg/dL, male sex was a statistically significant independent predictor of heart disease after controlling for other clinical variables. The findings support sex-specific screening and preventive strategies for high-cholesterol male patients and demonstrate the value of interpretable machine learning models for clinical decision support. Larger externally validated datasets are needed to assess generalizability.
Cardiovascular Diseases (CVDs) continue to be one of the leading causes of deaths in the world, claiming some 17.9 million lives every year. This burden is higher in Pakistan because of "Asian Indian Phenotype" which makes them vulnerable to early coronary artery disease. The commonly used traditional risk prediction models, including the Framingham Risk Score, have been developed in Western populations and are poorly predictive in South Asian populations. This study aims to fill this important gap by designing, implementing and comparative evaluation of six supervised machine learning algorithms for early detection of cardiovascular disease using a locally collected clinical dataset of 411 patient records with 13 independent clinical attributes. The models tested are Logistic Regression, K Nearest Neighbor, Support Vector Machine, Random Forest, Gradient Boosting and XGBoost. A rigorous gender-based mean imputation and Z-score normalization was done and split in 80/20 ratio. Empirical results show that the Random Forest classifier has Area under the Curve (AUC) of 0.9842, accuracy of 95.2%, precision of 96.0% and recall of 96.0%. The model was then exported and used to create a browser-based, predictive application that could be embedded in an interactive dashboard for real-time cardiovascular risk without the need for a server. These results confirm the effectiveness of ensemble learning approaches for medical diagnostics and highlight the potential of implementing ML-based screening tools in the limited resource healthcare environment in Pakistan.
Awais Khursheed, Soban Ahmed, Sibghat Ullah et al.· International Journal of Inn...· 0 citations
It is demonstrated that machine learning models can effectively identify high-risk patients and that minority data resampling significantly improves mortality classification reliability, and the approach offers potential value for clinical decision support systems and prioritised care pathways.
Abiodun Ojo· Journal International Review...· 0 citations
The findings suggest that Logistic Regression is more suitable as a decision-support model for early heart disease screening due to its higher sensitivity, accuracy, and specificity.
Yan Risa, Aspi Sururi, M. Asadullah et al.· International Journal for Sc...· 0 citations
Hypertension is one of the most important modifiable risk factors for Cardiovascular Disease (CVD), yet identifying which hypertensive patients are at higher risk remains challenging in clinical practice. This study developed and evaluated three machine-learning models: logistic regression, random forest, and Gradient Boosting for CVD risk prediction in a cohort of 23,543 hypertensive patients drawn from a 70,000 patient cardiovascular dataset. After preprocessing, feature engineering, SMOTE-based class balancing, and hyperparameter tuning via randomized search, model performance was assessed on a held-out test set and validated using 5-fold stratified cross-validation with SMOTE correctly nested inside each fold to avoid data leakage. On the test set, tuned Gradient Boosting model achieved the highest accuracy (78.59%) and AUC-ROC (0.6681), outperforming Logistic Regression (0.6633) and Random Forest (0.6508). cross-validation provided a slightly different perspective: Logistic Regression’s mean AUC-ROC (0.6628) edged out Gradient Boosting (0,6609) and Random Forest (0.6383), SHAP analysis on the Gradient Boosting model identified systolic blood pressure, age, and height as the strongest predictors, with height rivaling systolic blood pressure and surpassing BMI a notable difference from Random Forest’s feature importance ranking. Lifestyle factors (smoking, alcohol, physical activity) contributed minimally. These findings highlight blood pressure and body size measures as the dominant clinical signals in this dataset, while demonstrating the potential of an explainable machine-learning model based on routinely collected clinical data to support cardiovascular risk stratification and clinical decision-making in hypertensive patients, despite their moderate discriminative performance.
C. M. Anyanwu, J. C. Onyianta, Ogechi Gift Onyedi et al.· Nature Journal of Emerging S...· 0 citations
Diabetes mellitus is a chronic metabolic disorder whose global prevalence continues to rise, creating an urgent need for scalable, low-cost tools for early risk identification. This study evaluates the effectiveness of five machine learning algorithms (Random Forest, XGBoost, Support Vector Machine [SVM], CatBoost, and TabNet) for predicting diabetes risk from routinely available clinical and lifestyle variables. Using the Pima Indians Diabetes dataset, a preprocessing pipeline was applied that included median imputation of physiologically implausible zero values, standardized (Z-score) feature scaling, Boruta-based feature selection, and class-imbalance handling through class-weight adjustment and the Synthetic Minority Oversampling Technique (SMOTE). Models were trained on an 80/20 train-test split and assessed using accuracy, precision, recall, F1-score, and the area under the receiver operating characteristic curve (ROC-AUC). Random Forest achieved the strongest overall performance (accuracy 0.753; F1-score 0.689; ROC-AUC 0.810), followed by CatBoost, SVM, and XGBoost, whereas TabNet performed worst with very low recall for the diabetic class. The best-performing model (Random Forest) was deployed in a lightweight Flask web application that returns a probability-based diabetes risk assessment, categorising each prediction as low, moderate, or high risk together with a tailored recommendation. The findings confirm that ensemble tree-based methods, particularly Random Forest, provide a reliable, interpretable, and deployable basis for diabetes risk screening, especially in resource-constrained settings. Key limitations include dataset homogeneity, residual class imbalance, and limited feature coverage.
T. Olayinka· FUDMA Journal of Sciences· 0 citations
Purpose This study aimed to develop and validate an interpretable machine learning model to predict the 1- to 3-year risk of cardiovascular events in breast cancer patients by integrating baseline and treatment variables, while preliminarily investigating the potential association between short-term cardiac function decline and long-term adverse cardiovascular events. Methods We analyzed electronic medical records from 31,878 breast cancer patients. A composite cardiovascular event outcome was used. Predictors were selected via a two-step process: removing highly correlated variables (|r|≥0.7) and applying LASSO regression with 10-fold cross-validation, which refined 62 initial variables down to 18. Five models were built and compared using the area under the receiver operating characteristic curve (AUC-ROC). The optimal model was interpreted using SHapley Additive exPlanations (SHAP). Results Among 31,878 breast cancer patients, 3,960 (12.4%) experienced cardiovascular events. The XGBoost model demonstrated the best overall discriminative performance (AUC = 0.790). SHAP analysis identified endocrine therapy, anemia management therapy, and history of cerebrovascular disease as the top three predictors. Crucially, short-term decline in cardiac function was also selected as a significant predictor, supporting its role as a precursor to long-term events. Model robustness was confirmed via sensitivity analysis.
Luxin Wang, Rui Yan, Xinyu Zhu et al.· Frontiers in Oncology· 0 citations