Skip to content
#generative ai Preprint

Masking Is Not Enough: Generative Restoration for Multimodal De-Identification in Medical AI

Aug 2026 · 0 citations · 29 references
Computer Science

TL;DR

ClinX is introduced, an end-to-end multimodal PHI sanitization framework for medical image-text data, and results show that OCR-only masking is not sufficient as a standalone solution, and restoration-based sanitization better preserves clinically relevant visual context while sharply reducing recoverable PHI.

Abstract

Medical image-text data can expose protected health information (PHI) through both visible image content as well as accompanying text, creating a barrier to privacy-preserving medical AI systems. This risk is especially prominent in multimodal systems, where images, questions, reports, and clinical context may enter training, evaluation, or inference pipelines. Existing medical vision-language benchmarks primarily emphasize task utility, while de-identification methods are often evaluated separately from downstream reasoning. We introduce ClinX, an end-to-end multimodal PHI sanitization framework for medical image-text data. ClinX detects visible identifiers with optical character recognition (OCR), constructs binary PHI masks, and applies ClinX-PRISM, a no-skip generative restoration module with privacy-oriented post-processing for burned-in identifier suppression. In parallel, text-side PHI is reduced through progressive de-identification levels: regex masking, context-aware masking, and rewrite-based sanitization. We evaluate ClinX in medical visual question answering (MedVQA), jointly measuring PHI leakage and downstream utility across image-side, text-side, and combined de-identification settings. Results show that OCR-only masking is not sufficient as a standalone solution, and restoration-based sanitization better preserves clinically relevant visual context while sharply reducing recoverable PHI.

View source

Similar papers

Open access Jul 2026

Semantic de-identification of burned-in PHI in DICOM medical images: a deep learning–NLP pipeline validated on clinical and phantom TMM datasets

The growing adoption of AI-based healthcare research has increased the need for properly anonymized medical imaging datasets. PHI within DICOM files - particularly burned-in pixel-level text - poses significant privacy and regulatory risks. Existing methods either focus solely on metadata or remove all detected text indiscriminately, sacrificing clinically relevant annotations. This paper proposes a semantic de-identification pipeline integrating YOLOv11n-based text detection, domain-optimized EasyOCR, and a hybrid natural language processing (NLP) classification module combining regular expressions, keyword matching, and named entity recognition. A dual-path architecture processes metadata and pixel-level PHI in parallel, enabling complete DICOM sanitization while preserving non-PHI clinical annotations. The system was evaluated on 1,042 multi-modality DICOM images (CT, MRI, X-ray, ultrasound). As a secondary evaluation, the pipeline was also applied to two tissue-mimicking material (TMM) phantom datasets from TCIA - the RIDER Phantom MRI and Phantom FDA CT (RIDER = Reference Image Database to Evaluate Therapy Response; FDA = Food and Drug Administration) - which served as surrogates for controlled evaluation of metadata and burned-in identifier removal. The system achieves an F1-score of 95.4%, 96.1% recall, a structural similarity index measure (SSIM) of 0.969, a peak signal-to-noise ratio (PSNR) of 28.9 dB, and processes each image in 2.8 s. It achieves SSIM of 0.986 and PSNR of 49.0 dB on RIDER Phantom MRI, and SSIM of 0.974 and PSNR of 31.5 dB on Phantom FDA CT. These results confirm that the pipeline preserves quantitative pixel fidelity when applied to institutional and device identifiers embedded in phantom acquisitions, supporting blinding for domain-generalization studies across institutions. The modular design supports institutional customisation, making it suitable for clinical research workflows and privacy-compliant phantom imaging pipelines.

Remya Sethulekshmi, Manu J. Pillai, Nihal Ahammed et al. · 0 citations
#small language model Open access Aug 2026

Balancing privacy and performance: the impact of facial defacing on AI in medical imaging.

BACKGROUND Recent NIH Data Management and Sharing (DMS) policy updates and NIH controlled-access data security requirements have increased attention to facial anonymization and controlled-access handling of shared head imaging data. This is particularly relevant for datasets submitted to or hosted by the Cancer Imaging Archive (TCIA), where NCI Cancer Imaging Program/TCIA implementation practices address imaging data containing potentially reconstructable facial anatomy. While intended to protect patient privacy and strengthen public trust, defacing can distort craniofacial geometry and alter image statistics, potentially compromising the fidelity and reproducibility of artificial intelligence (AI) models trained on such data. Existing studies primarily validate visual anonymization quality, but few have quantified its downstream impact on deep learning-based medical imaging tasks. Understanding this privacy-utility trade-off is crucial for responsible data sharing and compliant AI development. METHODS We systematically evaluated three representative defacing algorithms, two invasive (QuickShear and Py-Deface) and one less destructive, facial replacement (mri_reface), across MRI and CT datasets from 600 subjects spanning three institutions. Model performance was assessed on three clinically relevant applications: (1) brain segmentation and Evans ratio biomarker quantification in normal pressure hydrocephalus (NPH) MRI using SLANT and FreeSurfer; (2) representative-slice selection and diagnostic reasoning for brain tumour MRI using vision-language models (VLMs); and (3) automated emergency head CT report generation using a fine-tuned Otter-based vision-language model. Each method's impact was quantified using Dice similarity, correlation metrics, reasoning accuracy, and natural-language generation scores (BLEU, METEOR, ROUGE, CIDEr). FINDINGS Invasive algorithms caused significant degradation across all tasks. QuickShear reduced mean Dice scores by up to 9% and introduced 14-19% failure rates during quality control, while PyDeface induced smaller but measurable performance losses. mri_reface maintained 100% success without any failures and achieved segmentation, diagnostic, and report-generation accuracy within 3-5% of the original data. Evans ratio distributions remained statistically consistent between mri_reface and original images (p > 0.05), whereas invasive methods introduced broader variance. Across all VLM tasks, mri_reface preserved high correlation with radiologist-selected slices (r = 0.979) and stable report-generation quality (BLEU-4 = 0.11 ± 0.06 vs. 0.12 ± 0.07 for original). INTERPRETATION Facial anonymization introduces a measurable privacy-utility trade-off that must be explicitly considered in the design of AI-ready medical imaging datasets. Invasive defacing compromises geometric and statistical integrity, reducing downstream model accuracy even outside facial regions. Facial replacement anonymization methods, such as mri_reface, effectively reconcile patient privacy with reproducibility, offering a practical path to NIH-compliant open data. Future regulatory and institutional policies should integrate quantitative privacy-utility assessment and mandate transparent reporting of anonymization pipelines to ensure that shared imaging data remain both ethically safe and scientifically valid under emerging digital health frameworks. FUNDING This work was partially supported by the American Heart Association (Award No. 25IPA1454088), the National Institutes of Health (Award No. 1R03CA286693-01A1 and Award No. 1R01CA291826-01A1), the U.S. Department of Defense (Award No. HT94252510807), and the National Science Foundation (Award No. 2545071).

Yuli Wang, Yuwei Dai, Haoyue Guan et al. · 0 citations
Preprint Aug 2026

MirrorNet: Can Medical Image Anonymization Really Protect Patient Identity?

Medical images are routinely de-identified---names, dates, and other metadata removed---and then shared for research, teaching, and public benchmarks under the assumption that this renders them anonymous. Such de-identification protects the metadata but not the pixels, and---apart from scans that directly contain facial structures---whether the image content itself identifies the patient has received little scrutiny. We investigate this question by learning a cycle-consistent correspondence between a cross-sectional medical image and a non-medical, patient-identifying image, using a pair of coupled, cycle-consistent variational autoencoders. From a held-out scan, the model recovers a recognisable likeness of the patient (identity-region MAE = 0.163); conversely, it synthesises a scan from such an image. These results indicate that a de-identified medical scan remains identifying---it is, in effect, a photograph of the patient---and that imaging data should be governed as biometric data rather than as anonymisable records. To support reproducibility, the code and trained models are shared at https://github.com/attilasimko/public-repository.

A. Simkó · 0 citations
Preprint Aug 2026

CRS-Bench: A Reference-Relative Reliability Benchmark for Medical Image Encoders

Pretrained image encoders are central to medical image classification, where expert annotation is costly and task-specific cohorts are often limited. As the model space expands from general-purpose to broad-medical and specialty-specific encoders, selecting the representation becomes a substantive modeling decision. Clean-test discrimination alone is insufficient for this purpose: encoders with similar AUROC can differ in calibration, label efficiency, and stability under acquisition perturbations or distribution shift. We introduce CRS-Bench, a controlled benchmark for multi-objective medical encoder selection. CRS-Bench evaluates 15 pretrained encoder families across dermatology, ophthalmology, and radiology using ISIC 2019, APTOS 2019, and CheXpert, with CheXpert-to-MIMIC-CXR as an observed institutional shift, yielding 17,575 controlled run records and 3,515 seed-aggregated metric rows. Each encoder is characterized along four operational reliability dimensions: discrimination, calibration, label efficiency, and robustness. We summarize these dimensions using the Clinical Reliability Score (CRS), a Pareto-aware, reference-relative score combining dominance, profile balance, and worst-axis performance. AUROC and CRS are positively associated but not decision-equivalent: 21 of 105 pairwise orderings reverse, with a mean absolute rank displacement of 1.87. Paired-seed bootstrap analysis identifies PanDerm, MedSigLIP, and MedGemma as a stable leading reliability tier rather than a statistically resolved single leader. CRS-Bench provides a controlled framework for selecting medical image encoders from multi-axis reliability profiles rather than clean-test AUROC alone.

Xingtao Lin, Hangqi Ren, Caiwan Sun et al. · 0 citations
Jul 2026

HiVLR: Hierarchical Vision-Language Reasoning for interpretable zero-shot radiography image understanding

Medical vision-language pre-training on image-report pairs has shown great potential to facilitate downstream image understanding tasks. However, prior approaches commonly exhibited limited accuracy on zero-shot image tasks, and lacked sufficient interpretability for unseen disease diagnosis, posing substantial usability concerns and trust issues for safety-critical medical applications. To alleviate them, we revisit how human doctors reason from a patient's radiology image for diagnosis, and propose a Hierarchical Vision-Language Reasoning (HiVLR) framework based on the clinical diagnostic workflow. In specific, we structure feature investigation into sequential rounds of thinking, i.e., (1) spotting the suspicious pathology observations (e.g., obscure) from all visual and textual inputs first and then (2) determining possible diagnostic findings (e.g., pneumonia) that match all pathology observations, to derive accurate disease predictions without compromising transparency in model decision making. Each round of thinking needs to analyze the inputted visual and textual embeddings by coarsely aligning them with cross-attention to establish global correspondences, highlight region-level visual features containing specific clinical content by prompt tuning-enabled fine-grained filtering, and then interpret the visual features in a condensed understanding to derive diagnostic-pertinent discoveries. Importantly, we enforce concept-level cross-modal compliance by ensuring that visual and textual features corresponding to the same clinical content are semantically consistent across concept dimensions (e.g., texture, shape, border). Based on this, we attach a concept-based interpretable diagnosis block to improve the accuracy and interpretability in downstream tasks simultaneously. Experiments showed that our approach greatly outperformed competing approaches on diverse zero-shot image tasks with superior interpretability.

Xilin Dang, Kang Li, Pheng-Ann Heng · 0 citations

Related blog posts