Skip to content
Open access

A medical image classification algorithm based on a hierarchical and complementary attention-enhanced Swin Transformer model

Aug 2026 · Scientific Reports · 0 citations

TL;DR

Comparisons across all datasets confirm that the proposed framework exhibits strong robustness and generalization capability when processing multiple medical imaging modalities, thereby providing reliable technical support for computer-aided medical diagnosis systems.

Abstract

With the rapid development of precision medicine and intelligent diagnostic technologies, automatic medical image classification has become an important tool for assisting clinical decision-making. However, substantial variations in lesion scale, complex long-range dependencies of tissue structures, and the need to capture subtle anatomical features present significant challenges to existing deep learning models. To address the limitations of the Swin Transformer in modeling multi-scale lesions and multi-level feature interactions, this study proposes a hierarchical complementary feature enhancement framework based on the Swin Transformer for medical image classification. The proposed architecture performs collaborative feature learning at three representation levels, including macro-scale lesion perception, global contextual interaction, and local detail refinement. Specifically, a Multi-scale Depthwise SE Block (MSD-SE Block) is introduced at the input of each stage of the Swin Transformer to enhance the model’s multi-scale feature representation capability. Subsequently, a Residual Convolutional Attention (RCA) module is integrated following the self-attention mechanism and the Multi-Layer Perceptron (MLP) to strengthen global contextual modeling, while a Local Detail Enhanced Residual Channel-Spatial Attention (LDERCSA) module is employed to refine subtle anatomical structures and discriminative local features. Through the coordinated interaction of these components, the proposed framework establishes a hierarchical feature enhancement mechanism that effectively improves medical image representation across multiple scales and feature levels. Comprehensive experiments were conducted on eight core subsets of MedMNIST v2, including BloodMNIST, BreastMNIST, DermaMNIST, OCTMNIST, OrganSMNIST, PathMNIST, PneumoniaMNIST, and RetinaMNIST, using an input resolution of 224 $$\times$$ 224. Single-module comparison and ablation studies demonstrate that each proposed component contributes positively to the overall performance. Experimental results show that the proposed model achieves significant performance improvements on BreastMNIST, OCTMNIST, OrganSMNIST, and PneumoniaMNIST. Furthermore, comparative evaluations across all datasets confirm that the proposed framework exhibits strong robustness and generalization capability when processing multiple medical imaging modalities, including ultrasound, CT, X-ray, endoscopic, and microscopic images, thereby providing reliable technical support for computer-aided medical diagnosis systems.

Read PDF

Similar papers

Open access Aug 2026

Hybrid CNN-Transformer Framework for Multi-Disease Detection from Medical Imaging Data

The proposed framework is intended to support clinical image assessment and prioritization rather than replace expert diagnosis, and demonstrates the potential of hybrid CNN-Transformer architectures for robust and scalable computer-assisted multi-disease screening from medical imaging data.

M. Balakrishnan, K. Ananthi, S. R. et al. · 0 citations
Open access Sep 2026

A Hybrid CNN–Transformer Framework for Enhanced Medical Image Classification

Medical image classification is important for computer-aided disease diagnosis due to its potential in identifying diseases from clinical images. Convolutional Neural Networks have been proved capable of modelling local spatial and textural features but are believed to be weak in capturing long-range dependencies. On t...

Shubham Vashishtha, Shiwangi Choudhary · 0 citations
Open access Sep 2026

MedFuse: dual-stream fusion of convolutional and vision transformer-based features for enhanced medical image classification

Medical image classification is fundamental to computer-aided diagnosis. Limited labeled samples, subtle inter-class differences, and heterogeneous lesion morphology make it difficult for a single representation to capture all relevant cues. This study investigates MedFuse, a simple dual-stream framework that combi...

Ya-Jing Ren, Hai Ling, Zheng Gu et al. · 0 citations
Open access Aug 2026

Multimodal Cancer Classification Using a SwinR Transformer Cross-Attention and Contrastive Learning

Deep learning models are designed to process complex medical images, allowing physicians to more accurately identify and classify cancerous cells. There precision and speed, are both crucial in real clinical settings. This paper presents an enhanced SwinR Transformer based framework that incorporates hierarchical featu...

Thirumurugan Shanmugam, B. Sowmiya, Mohamed AK Sadiq · 0 citations
Aug 2026

Automatic lung cancer classification using histopathological images: A hybrid DenseNet121-slantlet transform framework

Accurate classification of histopathological images is essential in the early diagnosis and treatment of lung cancer. Conventional deep learning approaches often face the challenges of simulating fine-grained texture features of low-level features as well as high-level semantic features that can be observed in complex...

Biswaranjan Debata, R. Priyadarshini, S. Mohapatra · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.