A Hybrid CNN–Transformer Framework for Enhanced Medical Image Classification
Abstract
Medical image classification is important for computer-aided disease diagnosis due to its potential in identifying diseases from clinical images. Convolutional Neural Networks have been proved capable of modelling local spatial and textural features but are believed to be weak in capturing long-range dependencies. On the other hand, Transformer architectures are good at capturing global contextual relationships through self-attention but consume lots of computation power and training data. This paper seeks to integrate these two approaches through developing an architecture for medical image classification. Specifically, this paper develops a CNN-Transformer hybrid architecture for medical image classification that fuses convolutional feature extraction and Transformer-based global context modelling. The developed architecture is tested using two publicly available medical imaging datasets covering different medical diagnosis problems. These datasets include HAM10000 dermatoscopic skin lesion images and chest X-ray pneumonia. Proper preprocessing and augmentation techniques are applied to increase the quality of images and generalization of models and strategies to handle class imbalance are used. The performance of the developed architecture is compared with that of CNN architectures (ResNet50, DenseNet121, and EfficientNet-B0) and Transformer architectures (ViT and Swin Transformer). Evaluation metrics include accuracy, precision, recall, F1-score, specificity and Area Under the Receiver Operating Characteristic Curve (AUC).