Evaluating the Predictive Performance of Metabolic Biomarkers for 20- Year Incidence Diabetes in the Yazd Healthy Heart Cohort: A Machine Learning Analysis
Sep 2026· Iranian Journal of Diabetes and Metabolism· 0 citations
TL;DR
Conventional machine learning models using baseline data are insufficient for predicting diabetes incidence over long-term horizons, and the development of dynamic models that incorporate temporal changes in variables and complex pathophysiological interactions is essential for acceptable accuracy in long-term prediction.
Abstract
Background: Type 2 diabetes is a major global health challenge, necessitating the development of accurate prediction models for early intervention. This study aimed to assess the performance of machine learning models in predicting the long-term incidence of diabetes and to identify predictive metabolic biomarkers based on data from the Yazd Healthy Heart Cohort.
Methods: This retrospective study was conducted on 906 non-diabetic individuals from the Yazd Healthy Heart Cohort with a 20-year follow-up. Five machine learning models (Logistic Regression, Linear Discriminant Analysis (LDA), Support Vector Machine (SVM), K-Nearest Neighbors (KNN), and Random Forest) were implemented based on 17 demographic, clinical, and biochemical variables. The data were randomly split into 70% for training and 30% for testing. Model performance was evaluated using 10-fold cross-validation and metrics including accuracy, sensitivity, specificity, F1-Score, and Area Under the Curve- Receiver Operator Characteristic (AUC-ROC).
Results: During the study, 340 participants (37.5%) developed diabetes. None of the models achieved satisfactory performance (AUC> 0.8). The best performance was observed for the LDA model in the 10-year prediction, with an accuracy of 73% and an AUC of 0.70, although its sensitivity was low (47%). The composite indices triglyceride-glucose (TyG) and atherogenic plasma index (AIP) were identified as the strongest predictors across most models.
Conclusion: Although composite biomarkers such as TyG show significant predictive potential, conventional machine learning models using baseline data are insufficient for predicting diabetes incidence over long-term horizons. The development of dynamic models that incorporate temporal changes in variables and complex pathophysiological interactions is essential to achieve acceptable accuracy in long-term prediction.
This study developed an AutoML-based CVD risk prediction model with superior discrimination and calibration, providing clinicians a practical tool for risk stratification, and has the potential to reduce CVD burden at the population level.
Si-Min He, Ju-Ping Wang, Le Zhao et al.· International Journal of Car...· 0 citations
Male sex was a statistically significant independent predictor of heart disease after controlling for other clinical variables and the findings support sex-specific screening and preventive strategies for high-cholesterol male patients and demonstrate the value of interpretable machine learning models for clinical deci...
T. Adeyemo· GSC Advanced Research and Re...· 0 citations
Evaluated machine learning algorithms for predicting diabetes risk from routinely available clinical and lifestyle variables confirm that ensemble tree-based methods, particularly Random Forest, provide a reliable, interpretable, and deployable basis for diabetes risk screening, especially in resource-constrained setti...
T. Olayinka· FUDMA Journal of Sciences· 0 citations
The results confirm the effectiveness of ensemble learning approaches for medical diagnostics and highlight the potential of implementing ML-based screening tools in the limited resource healthcare environment in Pakistan.
Awais Khursheed, Soban Ahmed, Sibghat Ullah et al.· International Journal of Inn...· 0 citations
ML models incorporating routinely measured liver enzymes improve T2D identification across US and Chinese datasets, indicating the potential utility of liver enzymes as accessible adjunctive indicators for diabetes risk stratification and improved T2D identification.
Ya-Ping Bi, Xiao-Jie Yuan, Yu-Feng Wang et al.· Frontiers in Endocrinology· 0 citations
To develop, compare, and externally evaluate machine learning (ML) and Cox regression models for predicting fasting plasma glucose (FPG)-defined incident prediabetes.
We performed a secondary analysis of a publicly available Chinese health-examination cohort and conducted an external comparative evaluation i...
Su Hu, Chang-Shun Yan, Xin-Yuan Tian et al.· Frontiers in Endocrinology· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.