Adversarial Evaluation of Multi-Organ Lassification in Computational Pathology with CNNS and Foundation Models
In this research, we aim to evaluate the robustness of convolutional neural networks (CNNs) and foundation models, such as ConCH and UNI, in classifying different types of organs from whole-slide images (WSIs) collected from various countries, different scanners, and clinical environments. Despite the dataset's inherent diversity, our results reveal that even a small perturbation (a white-box attack) with an intensity of 0.01 significantly impacts performance. ResNet-50 experienced a 67% reduction, ConCH 19%, and UNI 9.5% in accuracy. However, the foundation models performed much better, showing greater resilience and maintaining comparatively strong predictive performance even under the same intensity. Still, randomizing the dataset does not necessarily make your model resilient or robust, emphasizing the importance of thorough generalization and stability testing before clinical deployment. These models should be rigorously tested to evaluate their generalization and reliability prior to application in real patient settings. Foundation models look promising for building reliable and general AI-based cancer diagnostic systems, but they still need to prove their reliability in real-world clinical environments before large-scale adoption.