Jul 2026· 2026 International Conference on Intelligent and Sustainable AI Systems (ICOSAAS)· pp. 763-770· 0 citations· 20 references
Abstract
Diabetes has become a health problem worldwide. It often goes unnoticed until it causes health issues. Finding diabetes early using a lot of health and personal data can help reduce the diseases impact and healthcare costs. This study proposes a machine learning system for diabetes prediction. This system uses techniques to prepare data select important features handle unequal class distributions and combine multiple models. It is designed to process types of data from Electronic Health Records (EHRs) lifestyle factors and clinical measurements efficiently. Multiple machine learning models, for example tree-based classifiers, simple linear models and combined models are. Tested. Cross-validation is used to ensure the models are reliable and can be scaled up. The prediction of diabetes mellitus is based on identifying factors, so the importance analysis of characteristics is used to find the most influential predictors of diabetes. Oversampling of medical data involves the use of oversampling to overcome the problem of class distributions. The findings indicate that the given approach is more accurate, precise, possesses higher recall and F1-score, as well as ROC-AUC, compared to other models. This developed system offers an understandable solution for assessing diabetes risk early. It can be used in healthcare screening systems and clinical decision-support platforms for diabetes mellitus.
Experimental results demonstrate that Machine Learning techniques can effectively predict disease occurrence with high accuracy, thereby assisting healthcare professionals in early diagnosis and treatment planning.
Sunidhi, Mothe Rahul, M. Kumar et al.· International Journal for Re...· 0 citations
It is indicated that a rigorously conducted methodology and interpretability in machine learning development are crucial in creating machine learning solutions in healthcare decision support, which is the pathway to real applications in diabetes risk assessment.
T. Khan, M. Saeed, Majid Hussain et al.· Scientific Reports· 0 citations
Six supervised learning models were developed and compared for diabetes prediction using a dataset and compared for diabetes prediction using a 100k patients records with eight clinical features including gender, age, hypertension, smoking history, heart disease, BMI, HbA1c level, and blood glucose level.
A multi-model framework using wearable sensor data to improve early disease risk prediction and support timely clinical decisions enables more robust and interpretable disease risk predictions and supports health monitoring, early clinical interventions as well as evidence-based clinical and healthcare decisions.
K. Manivannan· Journal of Intelligent Decis...· 0 citations
Diabetes affects over 101 million people in India, with many more at risk due to routine and hereditary factors. Early diagnosis is crucial to prevent complications, which make accurate predictive tools essential in healthcare. This research uses Machine Learning (ML) algorithms to evaluate the likelihood of Type 2 Diabetes Mellitus (T2DM) using lifestyle and family history data. The trained models demonstrate strong predictive ability, allowing individuals to self-assess their risk and supporting healthcare professionals in early detection and intervention. This study presents a performance assessment of seven ML classifiers: Decision Tree (DT), Random Forest (RF), Support Vector Machine (SVM), Logistic Regression (LR), Naïve Bayes (NB), k-Nearest Neighbor (k-NN), and Extreme Gradient Boosting (XGBoost). These classifiers were applied to the widely used PIMA Indian Diabetes dataset (PIDD), which contains 768 clinical records of adult women aged 21 and above, providing key medical information for diabetes analysis. Multiple evaluation measures were applied to assess model performance with results showing that SVM achieved the highest accuracy and AUC, while LR, RF, and XGBoost also performed competitively. Although k-NN attained the highest recall, it yielded a higher false positive rate. These findings highlight that no single model is perfect for every situation, and the choice of classifier should match clinical needs. This study serves as a reference for ML applications in diabetes prediction.
Rizwan Akhtar, Muhammad Kalamuddin Ahamad· ITEGAM- Journal of Engineeri...· 0 citations
Diabetes is a prevalent long-lasting disease marked by high blood glucose due to inadequate insulin secretion and insulin resistance may result in serious life-threatening complications. Diabetes global prevalence has raised by fourfold over the past thirty years, and is the ninth foremost disease leads to death across the globe. Meanwhile, developments in machine learning presents new opportunities for prediction and classification of disease. However, despite of numerous existing models, a need for classifying types of diabetes still remains. The objective of this study is to develop an integrated, data-driven machine learning model for predicting the occurrence and classification of diabetes in an effort to improve on these limitations and investigate the ability to differentiate between types of diabetes through machine learning approach. To evaluate the proposed framework for classification of diabetes and its subtypes, publicly accessible diabetes-related datasets were employed for model development and evaluation. Binary classification is used for detecting occurrence of diabetes, while multiclass classification was employed for subtype classification, namely Prediabetes(PD), Type 1 Diabetes(T1D), Type 2 Diabetes(T2D), and Pancreatogenic (Type 3c-T3cD) Diabetes. K-Nearest Neighbors (KNN), Logistic Regression, Naive Bayes, Random Forest, and XGBoost machine learning algorithms were implemented. The XGBoost demonstrated the highest performance among all models with an accuracy of 0.97. Its feature importance scores validated predictive accuracy and to identify key factors that distinguish types of diabetes. The proposed model is intended to serve as a decision-support system for screening and classification tool using routine clinical data to facilitate early diagnosis and treatment planning.
T. Shobha, S. Pradeep, Seema Patil et al.· Scientific Reports· 0 citations