Aug 2026· Journal of King Saud University: Computer and Information Sciences· Vol 38· 0 citations· 71 references
TL;DR
A tri-stream interaction paradigm replacing symmetric skip connections with directional fusion among semantic, spatial, and decoder-propagated streams at each decoding stage, and a Global Prototype Bank that captures dataset-level anatomical regularities via attention-based retrieval and gated EMA updates, providing persistent semantic priors across images.
Abstract
Multi-scale feature fusion is a cornerstone of encoder-decoder architectures in medical image segmentation, yet effectively integrating representations across stages remains a significant challenge due to the inherent semantic–spatial gap. Deep features encode abstract semantic context but lack spatial precision, whereas early-stage features preserve fine-grained details but suffer from limited semantic discriminability. Existing fusion mechanisms, which often rely on symmetric aggregation or simple skip connections, fail to explicitly model the semantic-to-spatial guidance necessary for precise alignment. To address this, we propose a Tri-stream Prototype Fusion Network (TSPFusion) that introduces three key innovations: (i) a tri-stream interaction paradigm replacing symmetric skip connections with directional fusion among semantic, spatial, and decoder-propagated streams at each decoding stage; (ii) a Global Prototype Bank (GPB) that captures dataset-level anatomical regularities via attention-based retrieval and gated EMA updates, providing persistent semantic priors across images; and (iii) a Detail–Semantic Feature Aligner (DSFA) that performs semantic-guided refinement of spatial features prior to fusion, preventing feature interference from direct concatenation. Additionally, an Adaptive Pyramid Context Decoder module aggregates multi-scale information with resolution-aware dynamic pooling, and a Gradient-Gated Spatial Attention head enforces boundary-sensitive structural consistency. Extensive experiments on four medical imaging benchmarks (CT and Ultrasound) demonstrate that TSPFusion achieves state-of-the-art performance 97.82±0.85% DSC on COVID19 lung CT, 81.76% DSC on COVID19-Seg, and 88.16% mDice on cross-dataset BUSI→\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\rightarrow $$\end{document}STU, while maintaining a compact 5.92M parameter footprint.
Swin-DeepLabV3 is proposed, a hybrid semantic segmentation framework that integrates global and local feature modeling by combining a hierarchical Swin Transformer encoder with an Atrous Spatial Pyramid Pooling-based context module, demonstrating an effective balance between contextual representation and spatial precis...
Minh Cao Tran, Ha Minh Tan, Kien Cao-van et al.· IEEE Access· 0 citations
Polyp segmentation in colonoscopy images plays a pivotal role in computer-aided medical diagnosis and the early prevention of colorectal cancer. However, existing methods often suffer from performance degradation when confronted with extreme polyp scale variation and polyp boundary ambiguity. To address these challenge...
Tan Guo, Wen-Han Zhang, Fu-Lin Luo et al.· IEEE journal of biomedical a...· 0 citations
GLNet adopts a dual-branch encoder that combines a CNN-based Local Detail Perception Branch with a Mamba-based Global Context Modeling Branch, enabling the joint extraction of fine-grained local features and long-range semantic representations.
Medical image segmentation requires high accuracy and robustness, yet practical commercial deployment also demands privacy preservation and computational efficiency. In this context, the U-Net architecture, which can be inherently decoupled into independent encoder and decoder components, serves as a natural commercial...
Modern state-of-the-art deep learning architectures for medical image segmentation rely strictly on feed-forward passes over dense pixel/voxel grids, which scale poorly to large signals. Implicit neural representations (INRs) offer a lightweight, continuous alternative to raw grids, but are traditionally signal-specifi...
Kushal Vyas, Daniel Kim, T. Netherton et al.· Medical Image Analysis· 0 citations
Experiments demonstrate that the proposed cross-layer semantic alignment mechanism outperforms mainstream approaches in terms of Dice, HD95, and Intersection over Union metrics, validating the effectiveness of the cross-layer semantic alignment mechanism for complex medical image segmentation tasks.