Skip to content
Open access

Hybrid CNN-Transformer Framework for Multi-Disease Detection from Medical Imaging Data

Aug 2026 · International journal of computer information systems and industrial management applications · Vol 18, pp. 259-280 · 0 citations

TL;DR

The proposed framework is intended to support clinical image assessment and prioritization rather than replace expert diagnosis, and demonstrates the potential of hybrid CNN-Transformer architectures for robust and scalable computer-assisted multi-disease screening from medical imaging data.

Abstract

Medical imaging plays an essential role in the early detection and clinical assessment of multiple diseases; however, manual image interpretation is time-consuming and can be affected by inter-observer variability, particularly when subtle pathological patterns are present. Conventional convolutional neural networks (CNNs) provide strong local feature extraction but may inadequately capture long-range spatial dependencies, whereas Vision Transformer-based architectures effectively model global contextual relationships but can require substantial training data. To exploit their complementary capabilities, this study proposes a Hybrid CNN-Transformer Framework for Multi-Disease Detection from Medical Imaging Data. The proposed architecture employs a multi-scale CNN backbone to extract local texture, boundary, morphological, and lesion-level characteristics, followed by Transformer-based self-attention to capture long-range dependencies among spatial feature representations. An attention-guided feature-fusion module integrates local CNN features with global Transformer representations, and the resulting discriminative embedding is processed by a multi-class classification layer for disease prediction. Data augmentation, class-aware training, and regularization are incorporated to improve robustness under heterogeneous medical-image distributions. Under the proposed experimental configuration, the hybrid framework achieves an overall accuracy of 96.74%, sensitivity of 95.92%, specificity of 97.18%, precision of 96.31%, F1-score of 96.11%, and area under the receiver operating characteristic curve (AUC) of 0.986. Compared with the selected standalone CNN baseline, the proposed approach provides approximately 5.2% relative improvement in accuracy and 5.8% improvement in F1-score. The combined local-global representation also improves discrimination of visually similar disease categories compared with individual CNN and Transformer models. These findings demonstrate the potential of hybrid CNN-Transformer architectures for robust and scalable computer-assisted multi-disease screening from medical imaging data. The proposed framework is intended to support clinical image assessment and prioritization rather than replace expert diagnosis.

Read PDF

Similar papers

Open access Sep 2026

A Hybrid CNN–Transformer Framework for Enhanced Medical Image Classification

Medical image classification is important for computer-aided disease diagnosis due to its potential in identifying diseases from clinical images. Convolutional Neural Networks have been proved capable of modelling local spatial and textural features but are believed to be weak in capturing long-range dependencies. On t...

Shubham Vashishtha, Shiwangi Choudhary · 0 citations
Open access Aug 2026

A medical image classification algorithm based on a hierarchical and complementary attention-enhanced Swin Transformer model

Comparisons across all datasets confirm that the proposed framework exhibits strong robustness and generalization capability when processing multiple medical imaging modalities, thereby providing reliable technical support for computer-aided medical diagnosis systems.

Ya-Chao Si, Yi Zhang, Ming-Zhan Zhao · 0 citations
Aug 2026

Hybrid CNN–Transformer Architecture for Multi-Organ Disease Classification

Convolutional neural networks (CNNs) and Vision Transformers (ViTs) offer complementary strengths for medical image analysis: CNNs excel at capturing local texture and edge information through their inductive spatial bias, while transformers capture long-range dependencies through global self-attention but typically re...

S. Thanekar · 0 citations
Open access Aug 2026

TransCat: a hybrid CNN-transformer network with KAN for medical image segmentation

TransCat, a hybrid CNN-Transformer architecture for medical image segmentation, is proposed and an extended deformable attention mechanism with attentive value identification is developed, to control the computational burden caused by the enlarged token set.

Jin Wang, Zheng-Hua Yang, Dong-Ming Zhou et al. · 0 citations
Open access Aug 2026

Hybrid CNN-Transformer Based Brain Tumor Detection and Classification Using MRI: A Novel Framework with Multi-Level Attention and Regional Explainability

A novel hybrid convolutional neural networks and transformer architecture, HybCT-Net, augmented with a multi-level attention module and a regional explainability pipeline for brain tumor detection and classification is proposed, demonstrating superior performance than contemporary CNN, transformer and hybrid baselines.

Phanideep Karnati, Sukanya Roy, Dundi Urlamma et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.