Skip to content
Open access

Spearman-guided XGBoost Classification for Breast Cancer Diagnosis Using the Wisconsin Diagnostic Dataset

2026 · International Journal of Computer Theory and Engineering · 0 citations

Abstract

Breast cancer remains one of the leading causes of mortality among women worldwide, underscoring the critical need for effective and early diagnostic tools. This study presents a comprehensive Machine Learning (ML) framework that employs k-Nearest Neighbors (KNN), Random Forest (RF), Logistic Regression (LR), and Extreme Gradient Boosting (XGBoost) algorithms to analyze a breast cancer dataset obtained from the University of California, Irvine (UCI) repository. The dataset was partitioned into 70% training and 30% testing subsets to evaluate the generalization performance of each model. The Logistic Regression (LR) model achieved the highest accuracy at 96.5%, demonstrating its effectiveness in modeling linear decision boundaries. The RF algorithm followed with 95.3% accuracy, reflecting its capability in capturing complex, non-linear interactions. XGBoost achieved a competitive accuracy of 95.9%, highlighting its robustness in detecting subtle patterns and improving predictive performance. The KNN model recorded an accuracy of 92.9%, indicating its effectiveness in identifying localized patterns within the data. These findings underscore the potential of ML-driven approaches in enhancing breast cancer diagnostics and contribute meaningfully to public health by supporting timely and accurate disease detection.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.