Aug 2026· International journal of computer information systems and industrial management applications· Vol 18, pp. 259-280· 0 citations
TL;DR
The proposed framework is intended to support clinical image assessment and prioritization rather than replace expert diagnosis, and demonstrates the potential of hybrid CNN-Transformer architectures for robust and scalable computer-assisted multi-disease screening from medical imaging data.
Abstract
Medical imaging plays an essential role in the early detection and clinical assessment of multiple diseases; however, manual image interpretation is time-consuming and can be affected by inter-observer variability, particularly when subtle pathological patterns are present. Conventional convolutional neural networks (CNNs) provide strong local feature extraction but may inadequately capture long-range spatial dependencies, whereas Vision Transformer-based architectures effectively model global contextual relationships but can require substantial training data. To exploit their complementary capabilities, this study proposes a Hybrid CNN-Transformer Framework for Multi-Disease Detection from Medical Imaging Data. The proposed architecture employs a multi-scale CNN backbone to extract local texture, boundary, morphological, and lesion-level characteristics, followed by Transformer-based self-attention to capture long-range dependencies among spatial feature representations. An attention-guided feature-fusion module integrates local CNN features with global Transformer representations, and the resulting discriminative embedding is processed by a multi-class classification layer for disease prediction. Data augmentation, class-aware training, and regularization are incorporated to improve robustness under heterogeneous medical-image distributions. Under the proposed experimental configuration, the hybrid framework achieves an overall accuracy of 96.74%, sensitivity of 95.92%, specificity of 97.18%, precision of 96.31%, F1-score of 96.11%, and area under the receiver operating characteristic curve (AUC) of 0.986. Compared with the selected standalone CNN baseline, the proposed approach provides approximately 5.2% relative improvement in accuracy and 5.8% improvement in F1-score. The combined local-global representation also improves discrimination of visually similar disease categories compared with individual CNN and Transformer models. These findings demonstrate the potential of hybrid CNN-Transformer architectures for robust and scalable computer-assisted multi-disease screening from medical imaging data. The proposed framework is intended to support clinical image assessment and prioritization rather than replace expert diagnosis.
Medical image classification is important for computer-aided disease diagnosis due to its potential in identifying diseases from clinical images. Convolutional Neural Networks have been proved capable of modelling local spatial and textural features but are believed to be weak in capturing long-range dependencies. On t...
Shubham Vashishtha, Shiwangi Choudhary· International journal for ad...· 0 citations
Comparisons across all datasets confirm that the proposed framework exhibits strong robustness and generalization capability when processing multiple medical imaging modalities, thereby providing reliable technical support for computer-aided medical diagnosis systems.
Ya-Chao Si, Yi Zhang, Ming-Zhan Zhao· Scientific Reports· 0 citations
Convolutional neural networks (CNNs) and Vision Transformers (ViTs) offer complementary strengths for medical image analysis: CNNs excel at capturing local texture and edge information through their inductive spatial bias, while transformers capture long-range dependencies through global self-attention but typically re...
S. Thanekar· Natural Resources for Human...· 0 citations
TransCat, a hybrid CNN-Transformer architecture for medical image segmentation, is proposed and an extended deformable attention mechanism with attentive value identification is developed, to control the computational burden caused by the enlarged token set.
Jin Wang, Zheng-Hua Yang, Dong-Ming Zhou et al.· Frontiers in Bioinformatics· 0 citations
A novel hybrid convolutional neural networks and transformer architecture, HybCT-Net, augmented with a multi-level attention module and a regional explainability pipeline for brain tumor detection and classification is proposed, demonstrating superior performance than contemporary CNN, transformer and hybrid baselines.
Phanideep Karnati, Sukanya Roy, Dundi Urlamma et al.· International journal of com...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.