Jul 2026· International Conference on Intelligent Engineering Systems· pp. 649-654· 0 citations· 26 references
Abstract
Multi-site data collection enables the aggregation of large and diverse magnetic resonance imaging (MRI) datasets, which are essential for development of robust machine learning (ML) models in neuroimaging. However, site-related variability introduced by differences in scanner equipment and acquisition protocols (i.e. "batch effects") may confound downstream analyses and obscure meaningful information. Harmonization methods, aim to eliminate this site-induced variability from the data while preserving true biological signals through covariates integrated into the harmonization models. Beside harmonization, MRI quality control is also an essential step in data preparation. However, for multi-site data, image quality metrics (IQMs) designed to capture quality-related properties of the recordings, may also contain site-specific characteristics. Although harmonization methods such as ComBat, are widely used to mitigate batch effects, their impact on IQMs and the role of incorporated covariates remain insufficiently understood. To address these shortcomings, in this study, we evaluate the effects of different batch correction strategies on structural brain MRI IQMs by comparing simple data merging, database-wise standardization, ComBat without and with age and sex included as biological covariates through downstream application of ML models, and by statistical comparison of feature values for validation. We show that both database-wise scaling and harmonization reduce site-related information, however, nonlinear batch effects remain in the data. We also demonstrate that biological information is attenuated if not incorporated into the model as covariates, which in turn reduces harmonization effectiveness. Furthermore, we identify and analyze the most influential IQMs for site, age, and sex prediction across the different data processing strategies.
Accurate brain magnetic resonance imaging (MRI) segmentation remains challenging due to intensity inhomogeneity, acquisition-related bias fields, and ambiguous tissue boundaries. To address these challenges, a Multiplicative–Additive Bias Single-Function Dual-Level-Set (MAB-SFDLS) model is introduced within a Software-as-a-Service (SaaS)-based medical image analysis framework. The model incorporates both multiplicative and additive bias components into a unified variational energy formulation and employs a single level-set function with dual thresholds to achieve stable multi-region segmentation with smooth and continuous boundaries. The method was evaluated on the MRBrainS18 dataset, achieving Dice scores of 0.95 for white matter and 0.86 for gray matter, with a boundary deviation of 2.20 mm measured using HD95. Compared with the classical level-set formulation, notable improvements were observed in both overlap accuracy and boundary precision. The approach also demonstrated competitive performance against state-of-the-art deep learning models, including nnU-Net and U-Mamba, while maintaining lower computational requirements. Statistical analysis confirmed that the improvements were significant (p < 0.05). To enhance interpretability and practical applicability, the segmentation framework is integrated with a browser-based 3D visualization module that supports synchronized surface and volume rendering, as well as interactive region-of-interest exploration. This framework provides a practical, interpretable, computationally efficient, and scalable approach to robust brain MRI segmentation in a cloud-based medical imaging environment. The proposed model code and SaaS platform prototype are publicly available at
https://doi.org/10.5281/zenodo.20797546
.
Ala’a R. Al-Shamasneh, Amal Alshardan, Suad Alramouni et al.· Scientific Reports· 0 citations
Quality assurance (QA) in magnetic resonance (MR) imaging is critical but remains a challenging and time-intensive process, particularly when working with large-scale, multi-site imaging datasets. Manual QA methods are subjective, prone to inter-rater variability, and impractical for high-throughput workflows. Existing automated QA methods often lack generalizability to diverse datasets or fail to provide interpretable insights into the causes of poor image quality. To address these limitations, we introduce an unsupervised and interpretable QA framework for multi-contrast MR images that quantifies artifact severity. By assigning a numerical score to each image, our method enables objective, consistent evaluation of image quality and highlights specific levels of artifact presence that can impair downstream analysis. Our framework employs an unsupervised contrastive learning approach, leveraging simulated artifact transformations, including random bias, noise, anisotropy, and ghosting, to train the model without requiring manual labels or preprocessing. A margin-based contrastive loss further enables differentiation between varying levels of artifact severity. We validate our framework using simulated artifacts on a public dataset and real artifacts on a private clinical dataset, demonstrating its robustness and generalizability for automatic MR image QA. By efficiently evaluating image quality and identifying artifacts prior to data processing, our approach streamlines QA workflows and enhances the reliability of subsequent analyses in both research and clinical settings.
Savannah P. Hays, Lianrui Zuo, Blake E. Dewey et al.· Proceedings of machine learn...· 2 citations
Automated glioma segmentation in multi-modal magnetic resonance imaging (MRI) is critical in clinical neurooncology, yet it is challenged by large data volumes, tumor heterogeneity, and class imbalance. This study proposes an efficient and scalable system based on an optimized 2D U-Net architecture, integrated within a data engineering workflow utilizing HDF5. This integration enables processing large MRI datasets without loading the entire dataset into memory. The proposed method is evaluated on the public BraTS2020 benchmark using T1, T1ce, T2, and FLAIR modalities. The model achieved a Dice coefficient of 0.884, sensitivity of 0.851, specificity of 0.992, and an average Hausdorff distance of $\text{4. 2 ~ m m}$ on the test set. These results indicate segmentation accuracy consistent with expert annotations. Training was stable and converged within 15 epochs, attributed to the use of a Coefficient Dice loss function and batch normalization. Computationally, the 2D approach reduced the number of trainable parameters to approximately 7.8 million, allowing training and inference on consumer-grade GPUs with less than 8 GB of memory. The findings suggest that integrating an efficient model with optimized data management achieves a balance between segmentation accuracy and computational efficiency, making this approach suitable for resource-constrained clinical environments.
Lídices Reyes-Hung, Gabriel Trinke, I. Soto et al.· International Symposium on C...· 0 citations
PURPOSE
To evaluate the feasibility of a locally deployable large language model (LLM) system for automated MRI protocol selection addressing data privacy, annotation burden, and scalability limitations.
METHODS
This retrospective study included 598 German-language MRI order entries from three neuroradiology domains (brain, head/neck, spine) between June 2018 and January 2023. A radiologist labeled entries for 27 protocol classes based on institutional standard operating procedures (SOP). An SOP-grounded AI system using MedGemma 27B was developed to predict the MRI protocol from the order entry. The system was optimized using Stochastic Introspective Mini-Batch Ascent (SIMBA), a self-reflective prompt optimization algorithm, and compared with a hierarchical system that first classified the body region and then the MRI protocol. Data efficiency was evaluated using training subsets of 10-119 examples across 3 optimization runs per subset size.
RESULTS
The flat zero-shot model achieved 73.07% accuracy in the three-domain setting on the held-out dataset (n = 479). In the hierarchical model, prompt optimization yielded 73.90% ± 3.24% with 30 labeled examples and a maximum of 74.46% ± 3.82% but did not outperform the flat approach. Prompt optimization benefited the hierarchical system, whereas the flat model already performed strongly without labeled examples for optimization. Performance on brain and head/neck cases remained broadly stable after expansion from 22 to 27 protocol classes, i.e., including spine.
CONCLUSION
A locally deployable SOP-grounded open-weight LLM can support MRI protocol selection while preserving data privacy and needing minimal labeled data. In this dataset, hierarchical routing and prompt optimization did not improve overall performance over the flat baseline, although they altered optimization behavior and error profile. These findings support prospective evaluation in human-in-the-loop clinical workflows.
M. Vach, C. Boschenriedter, Daniel Weiss et al.· Clinical Neuroradiology· 0 citations
Whole-heart segmentation from CT and MRI is essential for quantitative cardiac image analysis, but remains challenging under multi-center and multi-modality distribution shift. In the CARE whole-heart segmentation task, models must generalize from limited labeled sites to unseen acquisition distributions, where variation in spacing, intensity, reconstruction texture, and anatomy can degrade out-of-distribution performance. We propose a modality-routed 3D cardiac segmentation pipeline that combines TotalSegmentator-initialized nnU-Netv2 models with site-characterized, label-preserving appearance augmentation. We first characterize the available sites using measurable image properties and use this analysis to motivate candidate data-space generalization routes. The final retained recipe applies Bias Field + Bezier appearance augmentation, combining smooth spatial intensity perturbation with nonlinear intensity remapping, followed by lightweight class-wise largest-connected-component cleanup. On the primary held-out-site validation splits, the final configuration improves CT mean Dice from 0.8350 to 0.9135 and MRI mean Dice from 0.7695 to 0.7830, while also reducing HD95. These results suggest that site-motivated appearance augmentation is a practical strategy for improving cross-site robustness in limited-data whole-heart segmentation. Our code can be found in https://github.com/Purdue-M2/Improving-Cross-Site-Whole-Heart-Segmentation
Tanishqua H Mudaliar, Justin Li, Daniel Lin et al.· 0 citations