Aug 2026· Ultrasound in Medicine and Biology· 0 citations· 39 references
Medicine
TL;DR
Multi-Scale CNN Token Transformer (MSCT-Trans), a lightweight and interpretable hybrid architecture for general-purpose ultrasound image classification, is proposed, which consistently outperformed CNN and Transformer baselines across accuracy, macro-F1 and area under the receiver operating characteristic curve, particularly under class imbalance and limited data regimens.
Abstract
Automated ultrasound image classification is increasingly important for clinical decision support in breast, thyroid and fetal screening. However, deploying deep learning models in such safety-critical settings demands not only high predictive accuracy but also transparency, interpretability and trustworthiness-properties that existing approaches address insufficiently. Convolutional neural networks (CNNs) capture local texture patterns but struggle with global contextual dependencies, while Transformer-based models offer long-range reasoning yet require large-scale training data and remain sensitive to ultrasound-specific noise, both limiting factors for clinical deployment. We propose Multi-Scale CNN Token Transformer (MSCT-Trans), a lightweight and interpretable hybrid architecture for general-purpose ultrasound image classification. MSCT-Trans extracts multi-scale feature maps from a pre-trained CNN backbone and converts them into a unified token sequence, enabling a Transformer encoder to model global dependencies and inter-scale interactions over semantically meaningful, noise-attenuated representations. To support clinical transparency, we conducted a two-part explainability analysis-Gradient-weighted Class Activation Mapping++ spatial localisation and softmax class probability breakdown-demonstrating that MSCT-Trans consistently attends to diagnostically relevant anatomical regions, produces well-calibrated confidence estimates and associates prediction errors with model uncertainty rather than over-confident mis-classification. Here we evaluated MSCT-Trans on three ultrasound benchmarks spanning breast (BUS-BRA + BUSI + UCLM), thyroid (TN5000) and fetal imaging. MSCT-Trans consistently outperformed CNN and Transformer baselines across accuracy, macro-F1 and area under the receiver operating characteristic curve, particularly under class imbalance and limited data regimens. The combination of strong predictive performance, spatially grounded interpretability and calibrated uncertainty estimation positions MSCT-Trans as a transparent and trustworthy foundation for ultrasound-based clinical decision support. Code: https://github.com/MohsinFurkh/MSCT-Trans.
TransCat, a hybrid CNN-Transformer architecture for medical image segmentation, is proposed and an extended deformable attention mechanism with attentive value identification is developed, to control the computational burden caused by the enlarged token set.
Jin Wang, Zheng-Hua Yang, Dong-Ming Zhou et al.· Frontiers in Bioinformatics· 0 citations
This work proposes MGT–UNet, a hybrid CNN–Transformer segmentation network that integrates multi-scale feature learning with global context modeling for breast ultrasound image segmentation and suggests that combining multi-scale feature learning with transformer-based global context modeling is beneficial for breast u...
H. Le, H. T. Huynh· BMC Medical Imaging· 0 citations
Over the years, Convolutional Neural Networks (CNNs) have demonstrated strong capability in cancer detection and classification using medical images. However, CNN-based models often struggle to capture long-range contextual dependencies. In such scenarios, integrating Compact Convolutional Transformer (CCT) architectur...
The proposed framework is intended to support clinical image assessment and prioritization rather than replace expert diagnosis, and demonstrates the potential of hybrid CNN-Transformer architectures for robust and scalable computer-assisted multi-disease screening from medical imaging data.
M. Balakrishnan, K. Ananthi, S. R. et al.· International journal of com...· 0 citations
Medical image segmentation requires high accuracy and robustness, yet practical commercial deployment also demands privacy preservation and computational efficiency. In this context, the U-Net architecture, which can be inherently decoupled into independent encoder and decoder components, serves as a natural commercial...
Comparisons across all datasets confirm that the proposed framework exhibits strong robustness and generalization capability when processing multiple medical imaging modalities, thereby providing reliable technical support for computer-aided medical diagnosis systems.
Ya-Chao Si, Yi Zhang, Ming-Zhan Zhao· Scientific Reports· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.