High-Dimensional Breast Cancer Classification: Evaluating Trade-offs Between Accuracy, Robustness, and Computational Cost
Abstract
Background: Breast cancer is the leading cause of cancer-related mortality in females, with 2.3 million new cases diagnosed annually. Machine learning (ML) algorithms have the potential to improve diagnostic accuracy through the analysis of high-dimensional data. However, the lack of standardized benchmarks across different algorithmic families impedes informed model selection for clinical deployment. Aim: This study addresses critical gaps in understanding the trade-off between classification performance and computational efficiency factors that are essential for implementing algorithms in resource-constrained healthcare settings. The research aims to compare the performance of different ML algorithms for breast cancer classification in a high-dimensional feature space. Patients and Methodology: The performance of 13 ML algorithms including conventional classifiers, ensemble methods, Gradient Boosting (GB), neural networks (NNs), and meta-ensembles was evaluated on the Database for Mastology Research with Infrared (DMR-IR) multimodal breast imaging dataset comprising 5,794 instances and 1,280 features. After excluding unconfirmed cases, the final analysis included 5,602 instances. A hybrid feature extraction pipeline combined deep features from three pre-trained convolutional neural networks (VGG16, ResNet50, DenseNet121) with handcrafted color, texture, and region-of-interest descriptors. Performance was assessed using accuracy, F1-score, and training time on a stratified 80:20 split with 5-fold cross-validation. Results: Ensemble methods significantly outperformed conventional classifiers. The Voting ensemble achieved the highest accuracy (99.62%) and F1-score (0.9962), followed by Stacking (99.58%) and Extreme Gradient Boosting (XGBoost) (99.58%). The Histogram Gradient Boosting (HGB) classifier demonstrated the optimal balance between accuracy (99.54%) and training time (5.98 seconds). Ensemble algorithms outperformed conventional classifiers by 9-10% in accuracy. Notably, the computationally efficient k-nearest neighbors (KNN) algorithm achieved 98.59% accuracy with only 0.02 seconds training time when applied to the hybrid feature representation. Conclusion: Ensemble methods demonstrate superior performance for high-dimensional breast cancer classification, with the Voting ensemble achieving maximum accuracy and HGB offering optimal practicality. Preliminary evidence-based model selection considerations are provided for different deployment contexts, from resource-constrained primary care to centralized reference centers.