Skip to content

Evaluating the Predictive Performance of Metabolic Biomarkers for 20- Year Incidence Diabetes in the Yazd Healthy Heart Cohort: A Machine Learning Analysis

Sep 2026 · Iranian Journal of Diabetes and Metabolism · 0 citations

TL;DR

Conventional machine learning models using baseline data are insufficient for predicting diabetes incidence over long-term horizons, and the development of dynamic models that incorporate temporal changes in variables and complex pathophysiological interactions is essential for acceptable accuracy in long-term prediction.

Abstract

Background: Type 2 diabetes is a major global health challenge, necessitating the development of accurate prediction models for early intervention. This study aimed to assess the performance of machine learning models in predicting the long-term incidence of diabetes and to identify predictive metabolic biomarkers based on data from the Yazd Healthy Heart Cohort. Methods: This retrospective study was conducted on 906 non-diabetic individuals from the Yazd Healthy Heart Cohort with a 20-year follow-up. Five machine learning models (Logistic Regression, Linear Discriminant Analysis (LDA), Support Vector Machine (SVM), K-Nearest Neighbors (KNN), and Random Forest) were implemented based on 17 demographic, clinical, and biochemical variables. The data were randomly split into 70% for training and 30% for testing. Model performance was evaluated using 10-fold cross-validation and metrics including accuracy, sensitivity, specificity, F1-Score, and Area Under the Curve- Receiver Operator Characteristic (AUC-ROC). Results: During the study, 340 participants (37.5%) developed diabetes. None of the models achieved satisfactory performance (AUC> 0.8). The best performance was observed for the LDA model in the 10-year prediction, with an accuracy of 73% and an AUC of 0.70, although its sensitivity was low (47%). The composite indices triglyceride-glucose (TyG) and atherogenic plasma index (AIP) were identified as the strongest predictors across most models. Conclusion: Although composite biomarkers such as TyG show significant predictive potential, conventional machine learning models using baseline data are insufficient for predicting diabetes incidence over long-term horizons. The development of dynamic models that incorporate temporal changes in variables and complex pathophysiological interactions is essential to achieve acceptable accuracy in long-term prediction.

View source

Similar papers

Sep 2026

Improving 10-year cardiovascular disease risk prediction using automated machine learning.

This study developed an AutoML-based CVD risk prediction model with superior discrimination and calibration, providing clinicians a practical tool for risk stratification, and has the potential to reduce CVD burden at the population level.

Si-Min He, Ju-Ping Wang, Le Zhao et al. · 0 citations
Open access Aug 2026

PREDICTING HEART DISEASE RISK FROM CLINICAL VARIABLES: A GENDER-SPECIFIC MACHINE LEARNING ANALYSIS AMONG HIGH-CHOLESTEROL PATIENTS

Male sex was a statistically significant independent predictor of heart disease after controlling for other clinical variables and the findings support sex-specific screening and preventive strategies for high-cholesterol male patients and demonstrate the value of interpretable machine learning models for clinical deci...

T. Adeyemo · 0 citations
Open access Aug 2026

A Comparative Evaluation of Machine Learning Algorithms for Diabetes Risk Prediction

Evaluated machine learning algorithms for predicting diabetes risk from routinely available clinical and lifestyle variables confirm that ensemble tree-based methods, particularly Random Forest, provide a reliable, interpretable, and deployable basis for diabetes risk screening, especially in resource-constrained setti...

T. Olayinka · 0 citations
Open access Aug 2026

Machine Learning-Based Early Cardiovascular Disease Prediction: A Comparative Analysis of Supervised Learning Algorithms Using a Pakistani Clinical Dataset

The results confirm the effectiveness of ensemble learning approaches for medical diagnostics and highlight the potential of implementing ML-based screening tools in the limited resource healthcare environment in Pakistan.

Awais Khursheed, Soban Ahmed, Sibghat Ullah et al. · 0 citations
Review Open access Sep 2026

Machine-learning-based prediction model of type 2 diabetes using liver enzymes: a cross-sectional study

ML models incorporating routinely measured liver enzymes improve T2D identification across US and Chinese datasets, indicating the potential utility of liver enzymes as accessible adjunctive indicators for diabetes risk stratification and improved T2D identification.

Ya-Ping Bi, Xiao-Jie Yuan, Yu-Feng Wang et al. · 0 citations
Open access Sep 2026

Development and comparative evaluation of machine learning algorithms and Cox regression for predicting fasting plasma glucose-defined incident prediabetes: a longitudinal cohort study

To develop, compare, and externally evaluate machine learning (ML) and Cox regression models for predicting fasting plasma glucose (FPG)-defined incident prediabetes. We performed a secondary analysis of a publicly available Chinese health-examination cohort and conducted an external comparative evaluation i...

Su Hu, Chang-Shun Yan, Xin-Yuan Tian et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.