A Hybrid CNN–Transformer Framework with Wavelet-Based Feature Extraction for Multimodal Cancer Detection
Finding cancer early and making a good treatment plan are both important for boosting survival rates. MRI, PET, and CT are advanced imaging techniques that have greatly improved cancer screening, staging, and therapy monitoring. However, their high costs and need for specialised equipment make them hard to get, especially in places with few resources. In this setting, optical imaging technologies are becoming affordable and portable options for finding cancer early. Image preprocessing techniques were used to improve data quality in order to deal with problems including dataset imbalance and noise. They used the Haar wavelet approach to extract features from medical photos that were important. A hybrid CNN-Transformer architecture was suggested, with four main parts: Shallow Feature Extraction (SFE), CNN/Transformer, Deep Feature Fusion (DFF), and up-sampling. This combination uses CNN to find local features and Transformers to find global dependencies, which makes deep feature learning strong and adaptable. The proposed model had an overall accuracy of 97.13%, a precision of 95.30%, a specificity of 96.74%, and an AUC of 98.74%. This shows that it is quite good at finding cancer. These results show that the model could provide accurate, quick, and easyto-use diagnostic solutions.