Skip to content
Conference Open access

A Novel Hybrid Deep Learning Architecture Combining CNNs, Vision Transformers, and Multi-Scale Attention for Enhanced Detection of Breast Cancer in Histopathology Images

2026 · ITM Web of Conferences · 0 citations · 15 references

TL;DR

The proposed HrybridViT-CAM, a hybrid deep learning system that integrates convolution neural networks, Vision Transformers, and multi-scale attention system in order to classify breast cancer using histopathology images was able to detect the malignant regions of interest (ROIs) like nuclei pleomorphism, atypia chromatin patterns and architectural distortions which are comparable to clinical diagnostic standards.

Abstract

Breast cancer has been one of the major causes of cancer mortality in the world. According to WHO report 2.5 million deaths are predicted as a result of breast cancer in the world in 2040. Although deep learning has demonstrated encouraging histopathology image analysis, the current methods frequently fail to provide local morphological information and global contextual information at the same time. In this paper, we present our HrybridViT-CAM, a hybrid deep learning system that integrates convolution neural networks( CNN ), Vision Transformers( ViT ) and multi-scale attention system in order to classify breast cancer using histopathology images. Its architecture has a two-way structure: a CNN arm (EfficientNetB7) to extract local features and a vision transformer arm to analyze the global context, fused together with a cross-attention fusion block. We use convolution block attention modules(CBAM) and deformal attention to improve feature discrimination. The model was tested on the BreaKHis dataset with various magnifications (40x, 100x, 200x, 400x) in binary as well as in multi-class classification. For binary classification (benign vs. malignant), HybridViT-CAM achieved accuracies of 99.87%, 99.42%, 98.95%, and 98.31% at 40χ, 100χ, 200x, and 400x magnifications, respectively. For eight-class subtype classification, the corresponding accuracies were 98.76%, 97.89%, 97.23%, and 96.54%, respectively. Grad-CAM++ and attention visualization techniques allowed explaining the results, which showed high correspondence(93.7% agreement) with pathologist diagnostic criterial. The proposed model was able to detect the malignant regions of interest (ROIs) like nuclei pleomorphism, atypia chromatin patterns and architectural distortions which are comparable to clinical diagnostic standards.

Read PDF

Similar papers

Sep 2026

An ensemble deep learning model for automated classification of breast cancer from histopathology images

A robust soft-voting ensemble-based deep learning model for automatic binary breast cancer identification using histopathology images can achieve effective classification performance without excessive attention complexity while keeping clear visual evidence.

M. Tiar, Nadjiba Terki, Z. Kahhoul et al. · 0 citations
Open access Sep 2026

Colorectal Lesion diagnosis using transformer and deep learning with multiscale feature interface

Colorectal cancer is the third most common malignancy worldwide. Manual screening requires expertise and resources. However, advancements in AI (artificial intelligence) have reduced the computation burden and time. Machine and deep learning have recently been used to diagnose colorectal lesions. The requirement of han...

D. P. Yadav, Bhisham Sharma, Julian L. Webber et al. · 0 citations
#artificial intelligence Preprint Sep 2026

A Lightweight CNN Integrated Compact Convolutional Transformer for Multi-Scale Feature Learning and reducing computational complexity for breast cancer mammography image detection and classification

Over the years, Convolutional Neural Networks (CNNs) have demonstrated strong capability in cancer detection and classification using medical images. However, CNN-based models often struggle to capture long-range contextual dependencies. In such scenarios, integrating Compact Convolutional Transformer (CCT) architectur...

M. Ahad, Ainuddin Ahmed · 0 citations
Conference Aug 2026

A Hybrid CNN–Transformer Framework with Wavelet-Based Feature Extraction for Multimodal Cancer Detection

Finding cancer early and making a good treatment plan are both important for boosting survival rates. MRI, PET, and CT are advanced imaging techniques that have greatly improved cancer screening, staging, and therapy monitoring. However, their high costs and need for specialised equipment make them hard to get, especia...

Deepika Upadhyay, S. Manikandan, R. V. S. Praveen et al. · 0 citations
Open access Sep 2026

Hybrid CNN–BiLSTM with Multi-Transformer Stacking for Skin Lesion Classification

A dual-branch deep learning framework that integrates a CNN–BiLSTM module for local spatial–sequential feature modeling with multiple transformer models (ViT, DeiT, SwinV2, SwinV2, and BEiT) for global representation learning is proposed.

Maryem Zahid, Mohammed Rziza, Rachid Alaoui · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.