Skip to content

Bound-Aware Per-Organ Recall Risk Control for Multi-Organ CT Segmentation under Clinical Domain Shift

Aug 2026 · 0 citations · 25 references
Computer Science

Abstract

Distribution-free risk control adds organ-specific recall guarantees to frozen segmentation. We calibrate per-organ thresholds for an AMOS-trained nnU-Net, audit transfer to RAOS, and estimate local re-certification cost using case-level voxel false-negative rate (FNR). The AMOS control passes, but $7/12$ organs exceed $\alpha{=}0.10$ after transfer; smaller calibration sets can mask exceedances with conservative or vacuous thresholds. Risk-Controlling Prediction Sets (RCPS) give high-probability control of population-mean risk, whereas Conformal Risk Control (CRC) gives weaker expectation control. Both require exchangeability; fixed and global thresholds give no per-organ guarantee. The Waudby--Smith--Ramdas (WSR) betting bound re-certifies six Tier-1 organs with 25 local cases, versus 30--40 for Hoeffding--Bentkus (HB). CRC needs 10--15 but has a heavier individual-case tail. No Tier-2 organ meets our illustrative precision criterion with 25 cases.

View source

Similar papers

Review Open access Aug 2026

Label-Free Threshold Selection for Out-of-Distribution Detection in Liver CT Segmentation

Reliable clinical deployment of automated liver segmentation requires mechanisms for detecting failures in rare and previously unseen scenarios. Achieving this goal requires an appropriately calibrated threshold that converts an out-of-distribution (OOD) score into a failure prediction. However, threshold calibration typically relies on expert-labeled failures, creating a substantial annotation burden when failures are rare. Building upon our prior work, which uses Pairwise Surface DSC scores as indicators of segmentation quality, we propose a label-free framework for calibrating OOD score thresholds. First, we fitted a log-t distribution to Pairwise Surface DSC scores from a validation set of 400 internal scans to approximate an in-distribution score distribution. New segmentations were assigned significance scores based on their extremity under this fitted distribution and categorized into Low, Medium, and High Risk review groups using statistically principled cutoffs of 0.25 and 0.05. The fitted log-t distribution provided a strong fit to the observed scores and remained robust to moderate contamination by OOD cases. On an independent test set of 500 internal and external scans, the combined Medium and High Risk categories achieved 100% sensitivity and 79% specificity, whereas the High Risk category alone achieved 78% sensitivity and 96% specificity. These results indicate that clinically meaningful failure detection can be derived from unlabeled data. Our code is available at https://github.com/marshalln7/Label_Free_OOD_Threshold_Selection.

M. Nielsen, A. Castelo, M. Altaie et al. · 0 citations
Open access Jul 2026

Organ Segmentation with Machine Learning Models

Accurate segmentation of abdominal organs in Computed Tomography (CT) underpins radiotherapy planning, surgical planning, and disease monitoring. Existing benchmarks rank architectures by a single aggregate Dice score, without per-organ statistical testing or boundary-sensitive metrics, even though models are chosen organ by organ for clinical use. We benchmark ten architectures spanning convolutional, attention-based, transformer, and state–space (Mamba) families on the AMOS CT dataset under one identical nnU-Net-style pipeline; we report per-organ Dice, 95-percentile Hausdorff Distance (HD95), and Normalised Surface Dice, with pairwise significance tested on an independent external dataset (TotalSegmentator). A competitive cluster of convolutional and Mamba models leads; rankings are stable on large organs but reshuffle by 10–13% on the small, geometrically complex ones, and boundary fidelity separates the models into tiers that the Dice ranking hides. This ordering largely holds on the external set (Spearman ρ=0.84). Selecting a model on aggregate Dice alone is therefore unsafe for organ-specific clinical tasks: per-organ overlap and boundary metrics should be the primary acceptance criteria for selecting a model before clinical deployment.

Alexandros Barmperis, Olga Menegaki, Anna Panagiotakopoulou et al. · 0 citations
Preprint Aug 2026

Parameter-Efficient pretrained-CT-to-MRI Transfer for Rectal Cancer Segmentation: Performance-Calibration Trade-offs

Accurate rectal cancer segmentation from magnetic resonance imaging (MRI) is essential for adaptive radiotherapy and tumor response assessment, but deployment also requires computational efficiency and informative, calibrated uncertainty estimates. We therefore introduce SWIFT, a SWin pretrained model wIth parameter-eFficient and Tumor-aware fine-tuning for rectal cancer segmentation. A Swin V2 encoder pretrained on 10,444 public 3D CT volumes using a DINOv2-style objective was adapted to T2-weighted MRI through four cumulative configurations: full fine-tuning (SWIFT), decoder compression (SWIFTe), low-rank adaptation (SWIFTe-LoRA), and a four-member LoRA-decoder ensemble (SWIFTe-LDE4). Geometric accuracy, tumor detection, radiomic agreement, and probability calibration were evaluated on a held-out 247-case test set from a single-institution cohort acquired using 1.5 or 3 Tesla GE scanners. Compared with SWIFT, SWIFTe reduced total parameters by 70.1% (from 72.8M to 21.8M) and increased tumor detection rate from 89.9% to 93.9%, while achieving a slightly lower median surface DSC (0.61 versus 0.62) and improved radiomic agreement. In a separate SWIFTe ablation, removing tumor-aware augmentation reduced detection from 93.9% to 89.9% but increased surface DSC from 0.61 to 0.64, demonstrating a detection-boundary-agreement trade-off. SWIFTe-LoRA used 14.6% of SWIFTe's trainable parameters while retaining similar segmentation performance. SWIFTe-LDE4 achieved the lowest calibration errors among the four configurations after temperature scaling (expected calibration error, 0.217; Brier score, 0.222), although the absolute expected calibration error indicates residual miscalibration. Similar efficiency-calibration patterns were observed using the public VoCo checkpoint, supporting robustness across pretrained initializations rather than external clinical generalizability.

A. Rangnekar, J. T. Gomez, J. Deasy et al. · 0 citations
Open access Aug 2026

Appearance-Aware Robustness Probes for Strict External Breast Ultrasound Segmentation under Multi-Source Joint Training.

BACKGROUND Breast ultrasound segmentation is sensitive to external-domain shift caused by scanner, acquisition, annotation and appearance variation. Shadowing, gain and contrast variation are particularly relevant given that ultrasound is not an optical imaging modality, yet many segmentation reports still rely mainly on randomly split or source-overlapping validation. METHODS We reformulated the study as a strict external validation analysis. Six model configurations-U-Net, CMU-Net, RTCMUNet, PMix, CGate and PMix+CGate-were evaluated under three source-composition protocols: all-source joint training, leave-BUS-UCLM-out training and leave-BUS-BRA-out training. BUS (n = 163) and BUSI-WHU (n = 927) were held out from training and validation under all protocols and used as fixed strict external out-of-distribution cohorts. The primary endpoint was image-level intersection over union (IoU), with paired image-level bootstrap (10,000 resamples) used to estimate 95% confidence intervals. A targeted source-inclusion sensitivity analysis additionally compared two training-pool compositions for RTCMUNet and PMix+CGate that differed only in whether a small set of previously excluded malignant-red lesion images from one source was added. RESULTS PMix+CGate under all-source joint training achieved the highest average strict external out-of-distribution IoU across the two held-out cohorts (equal-domain mean 72.19). Against RTCMUNet trained without BUS-UCLM, the cross-protocol delta was +1.32 percentage points; the 95% CI [-0.33, 2.98] crossed zero and did not support a cross-protocol advantage. In the source-inclusion sensitivity analysis, RTCMUNet and PMix+CGate responded in opposite directions to the same compositional change, producing a model-by-composition interaction in IoU (-4.31 percentage points, 95% CI [-6.66, -2.15]) whose image-sampling interval excluded zero. CONCLUSION PMix+CGate achieved the highest average strict external out-of-distribution IoU under all-source joint training, but its cross-protocol advantage over RTCMUNet was not supported by an interval that crossed zero. The source-inclusion analysis further indicates model- and cohort-dependent sensitivity to training-pool composition, rather than a universal benefit from adding data or a universally preferred configuration.

Lin Ma · 0 citations
Open access Jul 2026

DynU-Net: Dynamic Uncertainty-Aware Multi-task U-Net for Joint Lesion Segmentation and Classification in Medical Imaging

A Dynamic Uncertainty-aware Network (DynU-Net) is proposed, a multi-task framework that adaptively balances segmentation and classification through learnable per-task uncertainty parameters that consistently outperforms both single-task and existing multi-task baselines.

Ngoc Ly Tran, Thi Thu Thuy Nguyen, Ba-Hung Ngo et al. · 0 citations

Related blog posts