Skip to content
#software testing Review Open access

Clinical evaluation of novel deep learning‐based auto‐segmentation software: Utility and potential pitfalls

Aug 2026 · Journal of Applied Clinical Medical Physics · Vol 27 · 0 citations
Medicine

TL;DR

RatoGuide demonstrated favorable performance in typical cases, but accuracy declined in atypical cases with artifacts or altered anatomy, particularly for atypical cases and organs in high-dose gradient regions.

Abstract

Abstract Background Accurate contouring of target volumes and organs at risk is critical in radiotherapy. While deep learning (DL) models offer automated contouring, their clinical applicability to real‐world cases containing anatomical variations and artifacts requires rigorous validation. Purpose To evaluate the clinical accuracy and potential vulnerabilities of RatoGuide, novel DL‐based auto‐segmentation software, using a dataset including atypical cases derived from routine clinical practice. Methods This single‐center retrospective study included 69 thoracic and male pelvic cases. The cohort was intentionally selected to encompass diverse anatomies and artifacts (e.g., pacemakers, SpaceOAR implants, artificial femoral head replacements, and unilateral atelectasis). Auto‐contours generated by RatoGuide were compared with expert‐approved manual contours. Performance was evaluated quantitatively using the Dice Similarity Coefficient (DSC) and 95th percentile Hausdorff Distance (HD95), and qualitatively via a 5‐point visual assessment scale by four independent reviewers. Statistical comparisons between cohorts were performed using the Mann‐Whitney U test. Additionally, a dosimetric evaluation was conducted for male pelvic cases to assess clinical impact. Results In typical cases, the software maintained high segmentation accuracy (thorax: mean DSC 0.856, mean HD95 6.89 mm; male pelvis: mean DSC 0.874, mean HD95 4.11 mm). However, performance declined in atypical cohorts (thorax: mean DSC 0.808, p = 0.0457, mean HD95 12.22 mm, p = 0.0002; male pelvis: mean DSC 0.828, p = 0.1620, mean HD95 6.16 mm, p = 0.0075). Notable decreases in accuracy were observed in challenging scenarios, such as artificial femoral head replacements (DSC: 0.754) and unilateral atelectasis (DSC: 0.784). Qualitative assessment revealed that errors were primarily due to anatomical factors and artifacts. Furthermore, the dosimetric evaluation identified one critical false‐negative error where a dose constraint violation was overlooked when the DL contour was used. Conclusions RatoGuide demonstrated favorable performance in typical cases, but accuracy declined in atypical cases with artifacts or altered anatomy. For clinical implementation, rigorous visual verification and manual review by experts are essential, particularly for atypical cases and organs in high‐dose gradient regions.

Read PDF

Similar papers

Open access Aug 2026

Real-world clinical impact of implementing and updating a deep learning-based automatic contouring system in rectal cancer radiotherapy.

BACKGROUND Accurate delineation of target volumes and organs-at-risk (OARs) is a critical yet labor-intensive component of rectal cancer radiotherapy. While deep learning (DL)-based automatic contouring systems are increasingly used to address inter-observer variability and improve efficiency, artificial intelligence models require rigorous quality assurance and updates to reflect current technology. However, high-level evidence regarding the longitudinal real-world impact of implementing and iteratively updating these systems in clinical workflows is currently lacking. PURPOSE This study aimed to evaluate the real-world clinical impact of implementing and updating a DL-based automatic contouring system in rectal cancer radiotherapy to generate high-quality evidence of iterative updates. METHODS This longitudinal retrospective analysis included 150 patients divided into three cohorts: pre-implementation (n1 = 50), post-implementation (n2 = 50), and post-update (n3 = 50). Geometric similarities between unedited-automatic and final treatment contours were compared across cohorts. Failure rates were systematically analyzed. Six oncologists contoured 21 additional cases through manual, first-generation (Auto1), and second-generation (Auto2) system-assisted methods to evaluate contouring time, inter-observer consistency, and accuracy. Additionally, a 5-point Likert scale was used by two blinded senior oncologists to assess the clinical acceptability of the generated contours. RESULTS The mean Dice similarity coefficient (DSC) values of clinical target volume (CTV) before and after implementing the automatic contouring system were 0.87 ± 0.04 and 0.88 ± 0.04 (P = 0.067), while those of OARs were 0.80 ± 0.06 and 0.88 ± 0.05 (P < 0.001), respectively. Following the system update, they improved from 0.88 ± 0.04 to 0.93 ± 0.04 for CTV (P < 0.001) and from 0.88 ± 0.05 to 0.95 ± 0.02 for OARs (P < 0.001). The system update achieved an approximately 80.6% reduction in the mean failure rate. Auto2-assisted method decreased the total time by approximately 58.8% compared with the manual method, and 21.9% compared with the Auto1-assisted method. This method also demonstrated optimal inter-observer consistency (0.95 ± 0.03) and accuracy (0.94 ± 0.03) for CTV. In the blinded clinical evaluation, 99.2% (125/126) of the oncologist-revised final contours received a Likert score of ≥ 4, and Auto2-generated unedited contours showed significantly higher clinical acceptability than Auto1 (4.02 ± 0.25 vs. 3.26 ± 0.49, P < 0.001) CONCLUSIONS: Implementing an automatic contouring system provided crucial guidance for clinical practice. Its iterative update significantly reduced workload and inter-observer variation while enhancing contouring efficiency and quality.

Ningyu Wang, Tongzhen Xu, Yu-Jie Kang et al. · 0 citations
Review Jul 2026

Integration of site-specific deep learning models for automated OAR segmentation in craniospinal proton therapy: a geometric and dosimetric analysis.

PURPOSE This study evaluated the feasibility of integrating RayStation deep learning auto-segmentation (DLS) models-originally trained for adult head and neck (HN), thorax-abdomen (TA), and male pelvis (MP) regions-for automated organs-at-risk (OARs) delineation in craniospinal irradiation (CSI) with intensity-modulated-proton-therapy (IMPT), focusing on geometric accuracy, dosimetric impact, and clinical efficiency. METHODS Forty patients (aged 2-28 years) with CNS embryonal tumors were retrospectively analyzed. Their planning CT datasets were sequentially processed through the HN, TA, and MP DLS models to auto-segment 60 OARs. Thirty-one OARs were compared with expert contours using Dice Similarity Coefficient (DSC) and Hausdorff Distance (HD). Spearman's rank correlation was used to examine associations between geometric metrics and patient age, BMI, and craniospinal length. Dosimetric evaluation was performed by recalculating IMPT plans on DLS-derived OARs. RESULTS The mean auto-segmentation time was 4.9 ± 1.0 min per patient. Across 31 OARs, mean DSC and HD were 0.73 and 2.45 mm, with 41.9 % and 93.5 % meeting good geometric criteria (DSC > 0.8, HD < 4 mm). Geometric accuracy showed significant correlations (p < 0.05) with age, BMI, and craniospinal length for up to 11 OARs. Accuracy decreased in children (<12 years), especially for the thyroid, bowel, and pelvic structures. Mean and maximum dose deviations were within 105 and 190 cGy(RBE), except for the mandible, esophagus, and anorectum. CONCLUSION The integrated DLS framework achieved reliable geometric and dosimetric performance across most OARs with substantial efficiency gains, offering a practical solution for rapid, standardization of CSI planning workflows under expert review.

S. Gayen, D. Sharma, Gaganpreet Singh et al. · 0 citations
Open access Jul 2026

A prospective development and evaluation of a 2D convolutional neural network-based auto-segmentation model for cervical cancer radiotherapy

Accurate delineation of target volumes and organs at risk (OAR) is essential in radiotherapy planning for cervical cancer. Deep learning (DL)-based auto-segmentation has the potential to improve contouring efficiency and workflow. This study reports the prospective development and internal validation of a DL-based auto-segmentation model- Deep contour (DC) for cervical cancer radiotherapy. In this prospective single-institution study, a 2-dimensional convolutional neural network based on the LinkNet architecture, DC, was trained on 190 computed tomography (CT) datasets for abdominal and pelvic OARs and 90 cervical cancer datasets for target volumes. Independent validation was performed on 20 CT datasets. Model performance was evaluated using dice similarity coefficient (DSC), Jaccard Index (JI), 95th percentile Hausdorff distance (HD95), average symmetric surface distance (ASSD), and surface dice coefficient (NSD). Expert internal and external radiation oncologists qualitatively explored clinical acceptability using a Likert scale, and segmentation time was compared with manual contouring. The DC demonstrated the greatest geometric performance for the femur (DSC 0.92 ± 0.03; NSD 0.94 ± 0.04), bowel bag (DSC 0.89 ± 0.03; NSD 0.77 ± 0.12), and bladder (DSC 0.88 ± 0.14; NSD 0.87 ± 0.12). Among the target volumes, the inguinal nodal clinical target volume (CTVn_inguinal) achieved the greatest agreement (DSC 0.77 ± 0.04; NSD 0.76 ± 0.07). Moderate performance was observed for the rectum (DSC 0.75 ± 0.16), liver (DSC 0.73 ± 0.21), and pelvic nodal clinical target volume (CTVn_pelvis) (DSC 0.60 ± 0.10), whereas lower performance was observed for anatomically complex structures such as the duodenum, anal canal, common bile duct, pancreas, and pelvic vessels. Clinical evaluation of two cases revealed a Likert score of III-IV for key pelvic organs, such as the bladder, femur, pelvic bowel bag, rectum, and sigmoid. Auto-segmentation significantly reduced the segmentation time from 77 min to 5 s per dataset (p < 0.001). This prospective validation demonstrates that DC auto-segmentation model can achieve acceptable geometric performance congruent across multiple abdominal and pelvic OARs and reasonable geometric performance for the elective inguinal CTV volume. Further validation on larger datasets and evaluation of clinical workflow integration are warranted. CTRI, TRN: CTRI/2024/02/063055, Registration date: February 22, 2024.

S. Menon, Mahak Gupta, Aakriti Bhardwaj et al. · 0 citations
Review Aug 2026

Systematic Literature Review of Deep Learning Techniques for Lung Cancer Segmentation

This study provides a clear and actionable framework to bridge the gap between DL‐based segmentation research and clinical deployment, and proposes a roadmap for future research focusing on lightweight architectures, edge–cloud integration, federated learning and explainable AI (XAI).

Muhammad Sufyan, Jun Qian, Jianqiang Li et al. · 0 citations
Review Open access Jul 2026

Validation of open-source deep learning segmentation tools for automated glioma volumetry: a narrative review of Dice scores, workflow efficiency, and clinical RANO 2.0 implementation

Background Segmentation enables extraction of quantitative imaging features to enhance glioma diagnosis by volumetric measurements and treatment response assessment. This narrative review evaluates open-source software for glioma segmentation and alignment with Response Assessment in Neuro-Oncology (RANO 2.0) volumetric criteria. Approach In this narrative review, we evaluated thirteen open-source tools selected for multimodal MRI sequence support (T1W, T1CE, T2W, FLAIR), performance on public datasets (BraTS Challenge), and applicability to RANO 2.0 volumetry. Assessment included Dice scores, workflow efficiency, advantages, limitations, and clinical translation potential. Results Tools achieved Dice scores 0.73–0.92 for tumor subregions. Despite high analytical validation, clinical utility is limited: 85% of treatment response studies have bias risk in patient selection per QUADAS-2 appraisal. Critically, as of November 2024, no automated tools have been formally validated specifically against RANO 2.0 criteria, despite their 2023 emphasis on volumetric standardization. Conclusion Open-source segmentation tools show promise for standardizing glioma volumetry with emerging tools (GlioMODA, AutoRANO) explicitly targeting RANO 2.0-compatible volumetric assessment. Hybrid approaches combining open-source innovation with commercial clinical integration could optimize clinical translation.

M. Slachta, M. Halaj, K. Balazova et al. · 0 citations
Open access Aug 2026

Improved Deep Learning Segmentation of Pediatric Diffuse Midline Gliomas After Treatment.

PURPOSE To develop and validate a pediatric diffuse midline glioma (DMG) auto-segmentation tool optimized for longitudinal treatment response assessment across the disease course. MATERIALS AND METHODS In this multi-institutional retrospective study, we included patients aged 1-30 years with DMG from an institutional pediatric cancer center, BraTS-PEDs 2024, and PNOC007, a prospective trial of radiation followed by peptide vaccine plus poly-ICLC, and we trained nnU-Net-based DMGtracker using expert segmentations from 140 institutional pre- and post-treatment studies and all 261 BraTS-PEDs 2024 pre-treatment studies, using four-sequence multiparametric MRI (T1, T1 post-contrast, T2, and FLAIR). We externally validated the model on 88 annotated PNOC007 studies (n = 49 patients) and compared it with the BraTS-PEDs 2024 winning model using median Dice similarity coefficient (DSC) and relative volumetric difference (RVD) for whole-tumor and contrast-enhancing tumor segmentation using the Wilcoxon signed-rank test. RESULTS Training and internal testing used 153 scans (59 post-treatment) from 74 patients. Incorporating post-treatment data improved internal whole-tumor DSC for our trained model (0.94 [IQR 0.82-0.96] vs 0.93 [0.81-0.96]; p<0.001). On external validation, DMGtracker outperformed the BraTS-PEDs 2024 winning model for whole-tumor segmentation, with higher DSC (0.90 [0.72-0.95] vs 0.81 [0.66-0.90]) and lower RVD (9.6% [3.7%-31.6%] vs 16.8% [7.3%-39.2%]). This advantage was greatest in post-treatment scans (n = 50 scans, DSC 0.90 [0.73-0.94] vs 0.80 [0.58-0.88]; RVD 9.7% [3.8%-26.2%] vs 19.8% [12.6%-39.8%]; p<0.001 for both). In post-treatment scans, DMGtracker achieved clinically acceptable whole-tumor segmentation (DSC > 0.80) in 64.0% of cases, compared with 52.0% for the BraTS-PEDs winner. CONCLUSION Training DMG segmentation models with post-treatment scans substantially improves performance in longitudinal clinical trial imaging, enabling more accurate volumetric tracking and response assessment.

John Zielke, Francesca Romana Mussa, A. Zapaishchykova et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Aug 17, 2026

Q&A: Rethinking how innovation happens

In his latest book, Professor Eugene Fitzgerald examines the forces that turn breakthroughs into value — and why innovation resists simple formulas.