Skip to content
Open access

Clinical Biomarker-Based Prediction of Chronic Kidney Disease Using Explainable Machine Learning

Aug 2026 · Kidneys · Vol 15, pp. 47-57 · 0 citations

TL;DR

The results show how a combination of explainable ML and accessible clinical biomarkers can offer a precise, transparent, and clinically interpretable framework for early CKD diagnosis, risk stratification, and informed clinical decision making.

Abstract

Chronic kidney disease (CKD) is a progressive disease that needs to be diagnosed properly to slow the progression of the disease and its complications. The authors of this study suggest a clinical biomarker-based prediction framework that can be improved with explainable machine learning to enhance the accuracy and interpretability of the classification of CKD. Prior to the development of the models, the publicly available CKD dataset consisting of 400 patient records and 25 clinical attributes was preprocessed by imputing missing values, encoding categorical features, and normalizing the data. The performance of a range of supervised machine learning algorithms, namely Logistic Regression, Decision Tree, Random Forest, Support Vector Machine, K-Nearest Neighbors, Naïve Bayes and Extreme Gradient Boosting (XGBoost) was assessed using the standard performance measures. The best predictive model of the models evaluated was the Random Forest classifier. Explainable Artificial Intelligence (XAI) was used to incorporate the contribution of each biomarker and hemoglobin, serum creatinine, packed cell volume, specific gravity, and albumin were found to be the most significant biomarkers. The results show how a combination of explainable ML and accessible clinical biomarkers can offer a precise, transparent, and clinically interpretable framework for early CKD diagnosis, risk stratification, and informed clinical decision making.

Read PDF

Similar papers

Open access Jul 2026

A comparative and interpretable machine learning framework for reliable diabetes risk prediction.

It is indicated that a rigorously conducted methodology and interpretability in machine learning development are crucial in creating machine learning solutions in healthcare decision support, which is the pathway to real applications in diabetes risk assessment.

T. Khan, M. Saeed, Majid Hussain et al. · 0 citations
Open access Jul 2026

Diabetes Prediction System Using Machine Learning

Experimental results demonstrate that Machine Learning techniques can effectively predict disease occurrence with high accuracy, thereby assisting healthcare professionals in early diagnosis and treatment planning.

Sunidhi, Mothe Rahul, M. Kumar et al. · 0 citations
Open access Aug 2026

A Comparative Evaluation of Machine Learning Algorithms for Diabetes Risk Prediction

Diabetes mellitus is a chronic metabolic disorder whose global prevalence continues to rise, creating an urgent need for scalable, low-cost tools for early risk identification. This study evaluates the effectiveness of five machine learning algorithms (Random Forest, XGBoost, Support Vector Machine [SVM], CatBoost, and TabNet) for predicting diabetes risk from routinely available clinical and lifestyle variables. Using the Pima Indians Diabetes dataset, a preprocessing pipeline was applied that included median imputation of physiologically implausible zero values, standardized (Z-score) feature scaling, Boruta-based feature selection, and class-imbalance handling through class-weight adjustment and the Synthetic Minority Oversampling Technique (SMOTE). Models were trained on an 80/20 train-test split and assessed using accuracy, precision, recall, F1-score, and the area under the receiver operating characteristic curve (ROC-AUC). Random Forest achieved the strongest overall performance (accuracy 0.753; F1-score 0.689; ROC-AUC 0.810), followed by CatBoost, SVM, and XGBoost, whereas TabNet performed worst with very low recall for the diabetic class. The best-performing model (Random Forest) was deployed in a lightweight Flask web application that returns a probability-based diabetes risk assessment, categorising each prediction as low, moderate, or high risk together with a tailored recommendation. The findings confirm that ensemble tree-based methods, particularly Random Forest, provide a reliable, interpretable, and deployable basis for diabetes risk screening, especially in resource-constrained settings. Key limitations include dataset homogeneity, residual class imbalance, and limited feature coverage.

T. Olayinka · 0 citations
Open access Jul 2026

Predictive modeling of early diabetes diagnosis: An evaluation of XGBoost, support vector machine, and random forest classifiers

It is recommended that healthcare systems adopt XGBoost-based predictive models in clinical decision support tools for early screening, while future studies should validate these models using real-world clinical data to enhance reliability and generalizability.

Idehen Emmanuel Imafidon, Chikere Obinna Munachiso, Dominic Evans Onyebuchi et al. · 0 citations
Open access Aug 2026

Machine Learning-Based Early Cardiovascular Disease Prediction: A Comparative Analysis of Supervised Learning Algorithms Using a Pakistani Clinical Dataset

Cardiovascular Diseases (CVDs) continue to be one of the leading causes of deaths in the world, claiming some 17.9 million lives every year. This burden is higher in Pakistan because of "Asian Indian Phenotype" which makes them vulnerable to early coronary artery disease. The commonly used traditional risk prediction models, including the Framingham Risk Score, have been developed in Western populations and are poorly predictive in South Asian populations. This study aims to fill this important gap by designing, implementing and comparative evaluation of six supervised machine learning algorithms for early detection of cardiovascular disease using a locally collected clinical dataset of 411 patient records with 13 independent clinical attributes. The models tested are Logistic Regression, K Nearest Neighbor, Support Vector Machine, Random Forest, Gradient Boosting and XGBoost. A rigorous gender-based mean imputation and Z-score normalization was done and split in 80/20 ratio. Empirical results show that the Random Forest classifier has Area under the Curve (AUC) of 0.9842, accuracy of 95.2%, precision of 96.0% and recall of 96.0%. The model was then exported and used to create a browser-based, predictive application that could be embedded in an interactive dashboard for real-time cardiovascular risk without the need for a server. These results confirm the effectiveness of ensemble learning approaches for medical diagnostics and highlight the potential of implementing ML-based screening tools in the limited resource healthcare environment in Pakistan.

Awais Khursheed, Soban Ahmed, Sibghat Ullah et al. · 0 citations
Open access Jul 2026

A Modified Genetic Algorithm–Based Feature Optimization Framework for Cardiovascular Disease Risk Prediction

The paper presents the framework that mediates between optimization methods and predictive modeling to provide valuable information on the next generation of data-driven cardiovascular diagnostics to provide valuable information on the next generation of data-driven cardiovascular diagnostics.

Shamal Salunkhe, Smita Bharne, P. Patil et al. · 0 citations