Hybrid Mechanistic and Machine Learning Framework for Interpretable Cardiovascular Risk Prediction Across Public Cohorts
Abstract
Cardiovascular risk assessment remains a central challenge in both population-based prevention and high-risk clinical settings. We benchmarked a previously developed mechanistic model of cardiovascular ageing against machine-learning baselines across two distinct cohorts and evaluated both discrimination and probability calibration. The mechanistic model was based on an ordinary differential equation (ODE) framework and adapted to the observable interaction structure of two independent public datasets: the Framingham Heart Study cardiovascular risk dataset, with 10-year coronary heart disease as the outcome, and the Heart Failure Clinical Records dataset, with mortality as the outcome. ElasticNet logistic regression and XGBoost were evaluated as cohort-specific machine-learning benchmarks. Predictive performance was assessed using repeated stratified five-fold cross-validation with three repeats. In the Framingham cohort (n = 4240; 644 events), ElasticNet achieved ROC-AUC = 0.727 (95% CI 0.705–0.747), PR-AUC = 0.343, and Brier score = 0.116, while XGBoost achieved ROC-AUC = 0.715 (95% CI 0.694–0.737), PR-AUC = 0.325, and Brier score = 0.118. The observable mechanistic score showed weaker discrimination, with ROC-AUC = 0.547 and PR-AUC = 0.202. In the Heart Failure cohort (n = 299; 96 deaths), ElasticNet and XGBoost achieved ROC-AUC values of 0.770 and 0.771, respectively, whereas the mechanistic score achieved ROC-AUC = 0.534. Post hoc calibration improved probability scaling of the mechanistic score but did not restore discriminative performance. Overall, cohort-specific machine-learning models demonstrated stronger discrimination, while the external applicability of the mechanistic framework depended on the alignment of available predictors and clinical endpoints in the target cohort.