Skip to content
Open access

Group-Aware and Explainable Machine Learning Framework for Heart Disease Prediction Using Optimized CatBoost

Sep 2026 · International Journal of Innovative Science & Technology · 0 citations · 26 references

Abstract

Accurate and reliable heart disease prediction can support early risk assessment and improve clinical decision-making. However, repeated or duplicate observations in widely used heart disease datasets can introduce data leakage and lead to overly optimistic performance estimates. This study proposes a duplicate-aware and explainable machine learning (ML) framework for heart disease prediction. A combined dataset containing 1,328 records and 13 clinical features was systematically analyzed, resulting in 605 unique feature groups. To enable reliableA classification model which can identify a decision boundary that maximize the separation between the two outcome classes was added to SVM, one of the kernel-based classification models evaluation, the feature groups were partitioned into 484 development groups and 121 independent test groups. Six ML models, including Logistic Regression, Random Forest, XGBoost, LightGBM, CatBoost, and Support Vector Machine (SVM), were comparatively evaluated using five-fold group-aware stratified cross-validation. CatBoost achieved the highest mean ROC-AUC among the evaluated models and was subsequently optimized using the development data. The optimized CatBoost model achieved an accuracy of 88.64%, precision of 89.92%, recall of 87.22%, F1-score of 88.55%, and ROC-AUC of 93.32% on the independent test set containing 264 records. At the unique feature-group level, the model achieved an accuracy of 86.78%, precision of 90.91%, recall of 81.97%, F1-score of 86.21%, and ROC-AUC of 93.28%. SHAP-based explainability analysis identified chest pain type, thalassemia status, ST-segment slope, number of major vessels, and resting electrocardiographic characteristics as among the most influential clinical features. The findings demonstrate that combining duplicate-aware evaluation, optimized CatBoost modeling, independent group-level testing, and SHAP-based explainability can provide a more reliable, robust, and transparent framework for heart disease prediction.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.