Aug 2026· IEEE Transactions on Image Processing· Vol 35, pp. 9057-9070· 0 citations· 67 references
Computer ScienceMedicine
Abstract
Accurate medical image segmentation plays a vital role in clinical diagnostics by facilitating the precise delineation of anatomical structures and pathological regions. However, the performance of existing segmentation methods is often constrained by the scarcity of high-quality annotated datasets, as manual labeling is both labor-intensive and reliant on domain-specific expertise. To address this limitation without requiring additional annotations, we propose a novel multimodal segmentation framework that leverages medical text annotations as an auxiliary modality to complement visual information. In particular, our approach introduces a learnable encoding strategy for joint distribution modeling of image and text, which enables discriminative fusion and effectively suppresses cross-modal redundancy. Moreover, we innovatively design a frequency-domain prompt encoder based on the discrete wavelet transform (DWT) to capture multi-frequency features, thereby significantly enhancing the model’s ability to delineate fine-grained boundaries. Overall, our framework integrates cross-attention for effective cross-modal interaction, employs joint distribution modeling to enable discriminative and redundancy-reduced multimodal fusion, and incorporates auxiliary supervision to strengthen the learning of task-relevant features. Extensive experiments on nine public datasets across three clinical tasks—including cell, lung infection, and polyp segmentation—demonstrate that our method achieves competitive segmentation performance while maintaining favorable computational efficiency. Comprehensive ablation studies and feature distribution visualizations further validate the effectiveness and robustness of our proposed components. The code will be made publicly available at https://github.com/chenpeng052/MDFP
This work proposes an enhanced 3D segmentation framework, UAtten-Unetr, designed to improve segmentation accuracy and robustness in complex medical scenarios, and innovatively developed a unified loss function based on bimodal modality-specific Dice constraints and uncertainty regularization, optimized for synchronous...
MRD-UNet provides a practical balance between segmentation accuracy and computational efficiency and outperforms baseline CNNs and performs comparably to heavier transformer-based models while using significantly fewer parameters.
Musa Doğan, I. Ozkan· BMC Medical Imaging· 0 citations
In clinical practice, Magnetic Resonance Imaging (MRI) data frequently suffer from the absence of certain imaging modalities, which inevitably degrades predictive performance. Existing approaches typically treat different modalities as independent and non-interacting during modal feature extraction, despite the prese...
Precise medical image segmentation is essential to modern clinical workflows and biomedical research. However, current automated models often lack the flexibility, generalizability, and clinician control required to adapt to out-of-distribution data or novel classes without computationally expensive retraining. Further...
Paul Machauer, M. Reisert, Janis Keuper· IEEE Access· 0 citations
Semi-supervised medical image segmentation methods have drawn wide attention as they reduce reliance on heavily annotated data. However, existing models suffer from confirmation bias with limited annotations, and structural or parameter coupling hinders self-correction, especially for medical images with ambiguous boun...
Dong-Sheng Wang, Xiao-Han Lang· Biomedical engineering and p...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.