Skip to content

Diabetes Prediction Using Machine Learning Model: A comparative Approach

· 0 citations · 10 references

TL;DR

Six supervised learning models were developed and compared for diabetes prediction using a dataset and compared for diabetes prediction using a 100k patients records with eight clinical features including gender, age, hypertension, smoking history, heart disease, BMI, HbA1c level, and blood glucose level.

View source

Similar papers

Open access Jul 2026

A comparative and interpretable machine learning framework for reliable diabetes risk prediction.

It is indicated that a rigorously conducted methodology and interpretability in machine learning development are crucial in creating machine learning solutions in healthcare decision support, which is the pathway to real applications in diabetes risk assessment.

T. Khan, M. Saeed, Majid Hussain et al. · 0 citations
Conference Open access 2025

Examination of Diabetes Prediction Using Machine Learning

In the future, it is essential to establish a standardized validation framework, develop interpretable algorithms, integrate wearable non-invasive markers, and implement Bayesian racial modeling to promote early screening and personalized intervention, thereby revolutionizing the clinical prevention paradigm.

Dingnan Wu · 0 citations
Conference Jul 2026

A Large-Scale Machine Learning Framework for Early Diabetes Prediction

Diabetes has become a health problem worldwide. It often goes unnoticed until it causes health issues. Finding diabetes early using a lot of health and personal data can help reduce the diseases impact and healthcare costs. This study proposes a machine learning system for diabetes prediction. This system uses techniques to prepare data select important features handle unequal class distributions and combine multiple models. It is designed to process types of data from Electronic Health Records (EHRs) lifestyle factors and clinical measurements efficiently. Multiple machine learning models, for example tree-based classifiers, simple linear models and combined models are. Tested. Cross-validation is used to ensure the models are reliable and can be scaled up. The prediction of diabetes mellitus is based on identifying factors, so the importance analysis of characteristics is used to find the most influential predictors of diabetes. Oversampling of medical data involves the use of oversampling to overcome the problem of class distributions. The findings indicate that the given approach is more accurate, precise, possesses higher recall and F1-score, as well as ROC-AUC, compared to other models. This developed system offers an understandable solution for assessing diabetes risk early. It can be used in healthcare screening systems and clinical decision-support platforms for diabetes mellitus.

Thatikonda Krishna Kalyan Gupta, Oruganti Yashwanth Reddy, I. S et al. · 0 citations
Conference Jul 2026

Machine Learning-Based Cardiovascular Disease Prediction Model

Cardiovascular disease, as a highly prevalent chronic condition, has shown a continuously rising incidence in China and now ranks as the leading cause of death among both urban and rural residents. Mainstream Cardiovascular disease risk prediction models have mostly been developed based on European and American populations, which do not align well with the physical characteristics and disease patterns of the Chinese population. Moreover, traditional statistical methods have inherent limitations, further restricting the clinical applicability of these models. To address this, the present study constructed a Cardiovascular disease risk prediction model tailored to the Chinese population using machine learning algorithms based on the China Health and Retirement Longitudinal Study database. The dataset was split into a training set and a test set at a ratio of 7:3. Seven algorithms were employed for parallel modeling, and multi-dimensional performance comparisons were conducted. The study found that, in addition to traditional risk factors such as blood pressure and blood glucose, sleep indicators—including nap duration and nighttime sleep duration—were also important influencing factors for Cardiovascular disease. The comparative results demonstrated that the LightGBM model achieved the best predictive performance, with an AUC of 0.828, a recall of 0.717, and an F1 score of 0.568. The integration of SHapley Additive exPlanations further validated the internal logic and rationality of the model. This model can assist clinicians in risk assessment, thereby effectively improving the accuracy and efficiency of cardiovascular disease prediction.

Longfa Chu, Jing Lin, Zekai Li et al. · 0 citations
Open access Jul 2026

An interpretable machine learning framework for early-stage diabetes mellitus prediction using comparative classification models and SHAP

A machine learning-based framework enhanced with explainability is introduced, built around a structured data preparation process that handles categorical encoding, numerical scaling, and minority class oversampling through the SMOTE technique, positioning it as a trustworthy tool for assisting medical professionals in data-driven clinical decision-making.

N. J, Deekshitha U, K. V · 0 citations
Jul 2026

Prediction of Hypertension Using Machine Learning

Hypertension, commonly known as high blood pressure, is a major risk factor for cardiovascular diseases and premature mortality worldwide. Early detection and prevention are critical in reducing its health impact. This study explores the application of machine learning (ML) techniques to predict the likelihood of hypertension in individuals using clinical and demographic data. A variety of supervised learning algorithms, including Logistic Regression, Random Forest, Support Vector Machines, and Gradient Boosting, were evaluated for their predictive performance [1]. The dataset was preprocessed through feature selection, normalization, and handling of missing values to improve model accuracy.[2] Performance metrics such as accuracy, precision, recall, F1-score, and AUC-ROC were used to assess the models [4]. The results demonstrate that ML models can effectively identify individuals at high risk of hypertension, offering a valuable tool for early intervention and personalized healthcare [5]. This approach underscores the potential of artificial intelligence in supporting public health efforts and enhancing clinical decision-making. Key words: Logistic Regression, Random Forest, Support Vector Machines, and Gradient Boosting.

G. Vamsi, K. Bhargavi · 0 citations