Skip to content
Open access

Multimodal Cancer Classification Using a SwinR Transformer Cross-Attention and Contrastive Learning

Aug 2026 · Journal of Intelligent Decision Making and Information Science · 0 citations · 30 references

Abstract

Deep learning models are designed to process complex medical images, allowing physicians to more accurately identify and classify cancerous cells. There precision and speed, are both crucial in real clinical settings. This paper presents an enhanced SwinR Transformer based framework that incorporates hierarchical feature extraction with adaptive attention to maintain a higher degree of "awareness" of what is important. Instead of just one signal, we do multimodal fusion, meaning that we combine genetic data, radiology images, and histopathology slides into one and coherent view, albeit it's a bit different in each modality. We also add differential learning + self-supervised learning, which helps with generalization and resilience, in essence by enabling more efficient feature representation in the event that labelled datasets are limited. But to make it even more accurate we employ cross attention mechanisms, which we enable complementary information to be integrated dynamically across modalities. In addition, we use adaptive patch-based processing for local feature extraction as well as global – some patterns are subtle. End to end pipeline. The first is a modality-specific encoder that deals with images and genomics, and the next is “cross attention fusion.” Then we do a contrastive pretraining step and at the end a classification head yields the result. Comparing to other state-of-the-art designs such as Vision Transformers (ViT) and ConvNeXt and hybrid CNN transformer method, our accuracy, recall and interpretability give positive results. From the experimental results, the SwinR Transformer is clearly better than the CNN based models with the multimodal fusion and advanced learning techniques. That implies it may be a viable alternative for actual diagnostic uses in the fight against cancer. Compared with the existing methods, our method is quantitative with an accuracy of 91.3%, recall of 88.1%, and F1 score of 89.7% which is superior. Besides the modality specific evaluations, we also conduct ablation studies to quantify the contribution of each modality and each architectural component.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.