Skip to content

Improving 10-year cardiovascular disease risk prediction using automated machine learning.

Sep 2026 · International Journal of Cardiology · pp. 134759 · 0 citations · 39 references
Medicine

TL;DR

This study developed an AutoML-based CVD risk prediction model with superior discrimination and calibration, providing clinicians a practical tool for risk stratification, and has the potential to reduce CVD burden at the population level.

Abstract

Aims

To develop a cardiovascular disease (CVD) risk prediction model with improved accuracy and interpretability by integrating diverse risk factors and applying Automated Machine Learning (AutoML), thereby enhancing clinical utility over conventional models.

Methods

This is a prospective cohort study. Data were obtained from the Multi-Ethnic Study of Atherosclerosis (MESA), including baseline and fifth follow-up visits, comprising 4713 participants. Exercise and dietary data were harmonized via Metabolic Equivalent of Task (MET) and Healthy Eating Index-2015 (HEI-2015), respectively. Predictor selection was performed using the Boruta algorithm alongside Random Forest (RF) error rate cross-validation. Logistic regression, four traditional machine learning algorithms, and H2O AutoML were each applied for model training and evaluation. Finally, the best-performing model was further interpreted using SHapley Additive exPlanations (SHAP).

Results

A total of 21 predictors were selected, including age, sex, and Total Cholesterol (TC). Among the evaluated models, H2O AutoML outperformed other methods with an accuracy of 0.864, specificity of 0.892, precision of 0.610, F1 score of 0.670, and a Youden index of 0.635, achieving the highest AUC of 0.882 (0.846-0.918). SHAP analysis revealed the relative importance of predictors, with age, TC and Digit Symbol Score (DSS) ranking highest.

Conclusions

This study developed an AutoML-based CVD risk prediction model with superior discrimination and calibration, providing clinicians a practical tool for risk stratification. By enabling personalized prevention and early identification of high-risk individuals, this model has the potential to reduce CVD burden at the population level. Notably, DSS exhibited high importance and may represent a candidate risk marker.

View source

Similar papers

Sep 2026

Evaluating the Predictive Performance of Metabolic Biomarkers for 20- Year Incidence Diabetes in the Yazd Healthy Heart Cohort: A Machine Learning Analysis

Conventional machine learning models using baseline data are insufficient for predicting diabetes incidence over long-term horizons, and the development of dynamic models that incorporate temporal changes in variables and complex pathophysiological interactions is essential for acceptable accuracy in long-term predicti...

A. Ghadiri-anari, S. Namayandeh, M. Ahi · 0 citations
Open access Sep 2026

Non-Invasive Prediction of Metabolic Syndrome Using Explainable Machine Learning

Background: Metabolic syndrome (MetS) is a complex health problem significantly associated with cardiovascular diseases and type 2 diabetes mellitus. Traditional diagnostic approaches rely on invasive biochemical markers, which limit their accessibility. Here, we developed an explainable machine learning (ML) framework...

Islam A. Berdaweel, S. Al-Azzam, Ghaith M. Al-Taani et al. · 0 citations
Open access Aug 2026

Machine Learning- Based Cardiovascular Disease Risk Prediction in Hypertensive Patients: Explainable insights into Clinical Risk Factors

Findings highlight blood pressure and body size measures as the dominant clinical signals in this dataset, while demonstrating the potential of an explainable machine-learning model based on routinely collected clinical data to support cardiovascular risk stratification and clinical decision-making in hypertensive pati...

C. M. Anyanwu, J. C. Onyianta, Ogechi Gift Onyedi et al. · 0 citations
Open access Aug 2026

A Comparative Evaluation of Machine Learning Algorithms for Diabetes Risk Prediction

Evaluated machine learning algorithms for predicting diabetes risk from routinely available clinical and lifestyle variables confirm that ensemble tree-based methods, particularly Random Forest, provide a reliable, interpretable, and deployable basis for diabetes risk screening, especially in resource-constrained setti...

T. Olayinka · 0 citations
Open access Aug 2026

PREDICTING HEART DISEASE RISK FROM CLINICAL VARIABLES: A GENDER-SPECIFIC MACHINE LEARNING ANALYSIS AMONG HIGH-CHOLESTEROL PATIENTS

Male sex was a statistically significant independent predictor of heart disease after controlling for other clinical variables and the findings support sex-specific screening and preventive strategies for high-cholesterol male patients and demonstrate the value of interpretable machine learning models for clinical deci...

T. Adeyemo · 0 citations
Open access Aug 2026

Predicting treatment-related cardiovascular risks in breast cancer patients: development and validation of an interpretable machine learning model

Development and validate an interpretable machine learning model to predict the 1- to 3-year risk of cardiovascular events in breast cancer patients by integrating baseline and treatment variables and identified endocrine therapy, anemia management therapy, and history of cerebrovascular disease as the top three predic...

Lu-Xin Wang, Rui Yan, Xin-Yu Zhu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.