Skip to content
Preprint

Multimodal Assessment of Pancreatic Cancer Resectability Using Deep Learning

Jul 2026 · 0 citations · 29 references
Computer Science

TL;DR

A fully automated multimodal deep learning framework that jointly analyzes 3D contrast enhanced CT and structured clinical information to classify patients into the three National Comprehensive Cancer Network resectability categories (upfront resectable, borderline resectable, locally advanced).

Abstract

Accurate determination of pancreatic ductal adenocarcinoma (PDAC) resectability relies on evaluating how the tumor interacts with major peripancreatic vessels on CT imaging, yet expert assessment often shows substantial variability. We introduce a fully automated multimodal deep learning framework that jointly analyzes 3D contrast enhanced CT and structured clinical information to classify patients into the three National Comprehensive Cancer Network (NCCN) resectability categories (upfront resectable, borderline resectable, locally advanced). The approach uses a Swin-UNETR backbone to obtain anatomy aware image representations through auxiliary segmentation of pancreas, tumor, and vascular structures. These features are fused with a compact clinical embedding derived from 17 routinely collected variables and processed by a lightweight classification head. Model training is guided by a dynamic multitask objective that adapts the balance between segmentation and classification based on current tumor Dice performance, promoting feature representations that remain both anatomically informed and discriminative.

View source

Similar papers

Open access Jul 2026

Deep learning for preoperative MRI-based endometrial cancer staging prediction.

Endometrial carcinoma ranks among the most common malignancies of the female reproductive system. Accurate early-stage staging is essential for devising appropriate treatment plans and assessing patient prognosis. This study aims to enhance diagnostic precision by overcoming the limitations of traditional imaging methods and existing deep learning models.To address challenges such as dependency on physician expertise, inefficiency, and deficiencies in feature transmission, boundary detail restoration, and multi-scale feature integration, we propose a novel architecture termed GCMF-UNet (Group Convolution and Multi-Scale Fusion U-Net). Furthermore, for classification tasks, we introduce MSFA-Net (Multi-Scale Fusion Attention Network), which integrates a ResNet-18 backbone with a multi-scale feature aggregation module, squeeze-and-excitation (SE) attention, and a Swin Transformer for global contextual modeling. Experimental results indicate that GCMF-UNet improves Accuracy from 90.1% to 94.2% and Recall from 89.3% to 94.8% compared to the standard U-Net. In classification performance, MSFA-Net improves F1-score from 0.901 to 0.938 over baseline ResNet-18, demonstrating enhanced capability in identifying critical lesion features. The proposed GCMF-UNet and MSFA-Net architectures effectively mitigate limitations of conventional diagnostic and deep learning approaches, offering more accurate lesion segmentation and classification. These advancements offer a technical basis for further exploration of automated diagnosis and staging in endometrial carcinoma.

Caili Gong, Yetong Qi, Ying Su et al. · 0 citations
Open access Aug 2026

Predicting ki-67 expression in breast cancer via transformer and multiple instance learning on DCE-MRI

Accurate assessment of Ki-67 expression levels in breast cancer is crucial for determining prognosis and making informed treatment decisions. Current immunohistochemical methods relying on needle biopsy introduce sampling errors due to tumor spatial heterogeneity, making the development of non-invasive, precise preoperative prediction methods of significant clinical importance. This study aims to explore and compare advanced deep learning models based on dynamic contrast-enhanced magnetic resonance imaging (DCE-MRI) for noninvasive assessment of Ki-67 expression. This retrospective study analyzed preoperative DCE-MRI data from 308 patients with histologically confirmed breast cancer. Adjacent slices centered on the tumor’s most significant cross-section were obtained to create a 2.5-dimensional (2·5D) dataset. We innovatively developed two deep learning models using the same dataset (1): a Multi-Instance Learning (MIL) model that combines slice-level predictive features with Predictive Likelihood Histogram (PLH) and Bag-of-Words (BoW) techniques (2); a Transformer-based fusion model that directly captures global contextual relationships between slices via self-attention mechanisms. The predictive performance of both models was systematically compared with traditional radiomics and clinical models. On an independent test set, the Transformer fusion model demonstrated optimal predictive performance with an area under the curve (AUC) of 0.875, achieving accuracy, sensitivity, and specificity of 0.839, 0.848, and 0.833, respectively. The MIL model ranked second (AUC = 0.825), with both models significantly outperforming traditional radiomics models (AUC = 0.698) and clinical models (AUC = 0.648). Deep learning models based on 2·5D DCE-MRI, especially Transformer models that achieve global feature fusion through self-attention mechanisms, can effectively and non-invasively predict Ki-67 expression status in breast cancer, surpassing traditional methods. This model shows potential as a reliable tool to help clinicians accurately assess tumor proliferation activity before surgery.

Yiying Cao, Mi Lin, Yanshan Ouyang et al. · 0 citations
Conference Jul 2026

Enhanced Colon Cancer Detection and Classification with Deep Learning Techniques: A Comparative Analysis

Colon cancer represents a growing universal issue related to health, with prompt and accurate detection essential for indispensable to improving outcomes of people's health. Standard approaches, like colonoscopy and histopathology, while useful, tend to be invasive, time-intensive and prone to human interpretative bias. Recent developments in deep learning (DL) have facilitated the creation of automated systems that improve the precision, speed and uniformity of colon cancer diagnosis and categorisation. This paper offers a thorough comparative examination of various advanced DL models, including ResNet, DenseNet and MobileNet, applied to multi-modal imaging datasets consisting of colonoscopy images. Quantitative findings indicate exceptional accuracy, precision and recall in diagnostic tasks, with MobileNet DL models outperforming in tumour diagnosis and grading than other peer groups.

Sivakumar Rajendran · 0 citations
Open access Jul 2026

Automated Differentiation of Hepatic Cysts and Metastatic Tumors Using a Deep Learning Ensemble Framework on CT Imaging in Low-Resource Areas

This approach offers a practical and reliable way to classify hepatic lesions with minimal manual intervention and may help improve consistency in diagnosis and could be integrated into clinical workflows to support decision making in low-resource areas.

Mamoun Qjidaa, Amine Souadka, Anass Benfares et al. · 0 citations
Open access Aug 2026

Multimodal AI Methods for Predicting Pathologic Complete Response in Breast Cancer.

Neoadjuvant chemotherapy (NAC) can eliminate all invasive cancer in some breast cancer patients, achieving a pathologic complete response (pCR) that is associated with a better prognosis. Prediction of pCR from pre-treatment radiology imaging is challenging but could provide immense value in the treatment planning. We propose a novel multimodal deep learning approach that combines pre-treatment dynamic contrast-enhanced MRI and clinical data for predicting pCR before chemotherapy. It is the first to integrate attention-based multiple instance learning technique for slice aggregation and a self-supervised contrastive objective to align image and clinical embeddings. Our model was trained on 1491 patients across four cohorts with expert tumor segmentations. On a stratified hold-out test set, the multimodal model achieved AUC 0.83 (95% CI 0.78-0.89), outperforming late fusion without contrastive alignment and single-modality baselines. These results, obtained on the largest multi-cohort breast MRI dataset comprising four clinical trials, underscore the potential of contrastive learning-based multimodal AI models to improve pCR prediction. With further validation, our proposed approach could support pre-treatment decision-making by identifying patients likely to achieve pCR and those who may benefit from alternative regimens.

Manu Goyal, Tanmay Shukla, Saeed Hassanpour · 0 citations