Development of Hybrid Convolutional Transformer Networks for Automated Segmentation of Volumetric Medical Imagery
Abstract
The automatic segmentation of volumetric medical images is an essential task in computer-aided diagnosis, clinical decision support, treatment planning and quantitative biomedical analysis. Three-dimensional medical structure segmentation, however, is still difficult due to low contrast, intensity inhomogeneity, complicated anatomical boundaries, noise, inter-slice differences, and large variations in organ/lesion morphology. Conventional convolutional neural networks (CNNs) are good at capturing local spatial features and cannot model long-range dependencies over large volumetric regions. Transformer-based models, on the other hand, offer excellent global context modeling capabilities and can consume significant computational power but might not necessarily excel at capturing fine-grained details. To overcome these limitations, this paper introduces a novel Hybrid Convolutional-Transformer Network (HCT-Net) for automatic segmentation of volumetric medical imagery. The proposed framework is designed to fuse multi-scale features and attention mechanisms to combine hierarchical convolutional feature extraction with transformer-based global representation learning. The preprocessing module normalizes intensities, resamples the volume, removes noise, aligns the structure in space and the hybrid encoder captures local anatomical structures and long-range contextual relationships simultaneously. A skip-connected decoder can be used to progressively reconstruct high-resolution segmentation maps with boundary information. A composite loss function, comprising of Dice loss and cross-entropy loss, is employed to tackle class imbalance and enhance the segmentation accuracy of the model. The performance is tested using Dice Similarity Coefficient, Intersection over Union (IoU), Pixel Accuracy, Precision, Recall, Hausdorff Distance, and inference time.