Leveraging Ancestral Health Records for Early Prediction of Cardiovascular Risk in Future Generations
Abstract
Cardiovascular disease (CVD) is the leading global cause of morbidity and mortality, with onset driven by a complex interplay between genetic susceptibility, demographic characteristics, and modifiable lifestyle factors. Most current risk-assessment systems rely heavily on just a patient’s existing clinical markers—cholesterol, blood pressure, body-mass index, and lifestyle factors—while largely neglecting the inherited and hereditary features that clearly predispose people across generations to cardiovascular disease. It is this gap that this study fills by proposing a scalable, distributed predictive framework that combines multi-generational ancestral health history features with demographic, clinical, and lifestyle data to predict an individual’s risk of future cardiovascular events before the appearance of clinical symptoms. To account for volume, variety and sparsity of cross-generational records, the integrated genealogical health corpus is stored in the Hadoop Distributed File System (HDFS) and processed there. The diLDA, a Distributed Latent Dirichlet Process, mines latent hereditary “disease-trend” topics from ancestral event histories; the DiNMF, a Distributed Non-negative Matrix Factorization, extracts interpretable low-rank latent risk factors from the high-dimensional patient–feature matrix. The two complementary latent representations are then fused with the raw clinical features and provided to a supervised classifier. Experiments show that on a three-generation cohort of 240,000 records synthesized from public cardiovascular datasets using simulated pedigree linkages, the proposed model achieves an accuracy of 95.25%, F1-score of 93.77%, and AUC of 0.969, far outperforming standard solutions including logistic regression, support-vector machines, random forests, XGBoost, and single-method latent baselines. The distributed pipeline also shows close-to-linear speed-up across cluster nodes. The main contributions are a feature representation that adapts to the ancestry of each user, a dual-factorization architecture that is distributed and incorporates this knowledge naturally into its input space, and empirical evidence for our hypothesis: that including ancestral information results in measurable predictive uplift. Incorporation of genomic markers and longitudinal validation will lead to future work that promotes precision-medicine-oriented preventive cardiology.