A lightweight Vision Transformer UNet is proposed that combines the hierarchical feature extraction capability of UNet with the global context modeling of Vision Transformers, enabling effective learning of both local and global features while maintaining computational efficiency with only 2.6 million trainable parameters.
Abstract
Accurate brain tumor segmentation from Magnetic Resonance Imaging is essential for diagnosis, treatment planning, and surgical guidance. Although Convolutional Neural Networks, particularly UNet, have achieved significant success in medical image segmentation, they often struggle to capture the long-range spatial dependencies required to model tumors with irregular shapes and complex boundaries. This paper proposes a lightweight Vision Transformer UNet that combines the hierarchical feature extraction capability of UNet with the global context modeling of Vision Transformers. The proposed architecture incorporates a compact ViT bottleneck within a U-Net encoder-decoder framework, enabling effective learning of both local and global features while maintaining computational efficiency with only 2.6 million trainable parameters. The model was evaluated on the TCGA LGG MRI Segmentation dataset, achieving a mean Intersection over Union of 0.8100 and a Dice score of 0.8446, outperforming the baseline UNet by 3.75% and 3.15%, respectively. Extensive quantitative and qualitative analyses, including confusion matrix evaluation, precision recall curves, per-image performance distribution, and tumor size dependency analysis, demonstrate the effectiveness and robustness of the proposed method for brain tumor segmentation.
The combined segmentation and classification results indicate that the proposed framework provides an effective approach for tumor localization and tumor-type identification, with potential applicability in automated brain MRI analysis and clinical decision support.
Malathi Janapati, Shaheda Akthar· International journal of re...· 0 citations
MMA-UNet, a novel hybrid architecture that integrates multiple complementary mechanisms within an encoder-decoder framework, demonstrates strong robustness on T1-weighted contrast-enhanced MRI, and its generalizability to larger datasets remains to be established.
Correct segmentation of brain tumors using multimodal magnetic resonance imaging (MRI) is critical for quantitative tumor evaluation, therapy planning and disease monitoring. However, the variability of tumor size, shape, location and appearance among MRI modalities poses a challenge for automatic segmentation. Heterog...
B. Omarov, D. Sultan, D. S. Rakhymberdiev et al.· NEWS of National Academy of...· 0 citations
The study demonstrates the potential of combining local feature extraction and global contextual learning to achieve more accurate and robust brain tumor segmentation from multimodal MRI images.
Lovedeep Kaur, Parminder Singh, Naveen Dhillon· International Journal of Com...· 0 citations
Accurate brain tumor segmentation from magnetic resonance imaging (MRI) is essential for diagnosis, treatment planning, surgical guidance, and disease monitoring. However, developing automated segmentation models that generalize across diverse tumor characteristics, imaging protocols, acquisition sites, and patient pop...
Mohammad Mahdi Danesh Pajouh, Sara Saeedi· 0 citations
Manual delineation of brain tumors in magnetic resonance imaging (MRI) is time-consuming and subject to inter-reader variability. Automated segmentation can support reproducible tumor delineation, but slice-wise processing does not explicitly represent contextual relationships among adjacent MRI slices. This study inve...
K. Santhi, S. Krishna· International Research Journ...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.