Predictive Analysis of Chronic Kidney Disease using Explainable Ensemble Learning Models
Abstract
Chronic kidney disease (CKD) progresses and is potentially fatal. Early detection of CKD is essential to prevent severe effects and kidney failure. This work provides an intelligent and comprehensible method of predicting CKD using ML and ensemble methods. It works with the CKD dataset from the UCI ML Repository, comprising 400 patient records from Apollo Hospital in India and containing 25 numeric and nominal attributes. We apply statistical and model-based imputation to replace missing values, and SMOTE to address class imbalance (CKD vs. non-CKD). We identified six clinically significant items, including haemoglobin, serum creatinine, albumin, high blood pressure, age, and diabetes mellitus. A combination of preprocessing, feature engineering, and feature selection techniques, such as Mutual Information, Variance Thresholding, RFE, and Sequential Feature Selection, was used to identify these. Various models are evaluated using Stratified K-Fold cross-validation. These are LR, Gaussian NB, SVM, DT, RF and AdaBoost. A mixed Voting Classifier composed of RF and AdaBoost achieves an F1-score of 99.3% and 99.4% accuracy, precision, and recall. The model is easier to comprehend because individual forecasts and feature contributions are explained using LIME and SHAP. The web application, developed using Flask, allows users to register and log in, enter data securely, make real-time predictions, and receive interpretable results. It provides outcomes such as “CKD detected” and “no CKD detected” that help doctors make informed decisions