Aug 2026· International Journal of Medical Informatics· Vol 221, pp.
106672
· 0 citations· 41 references
Medicine
TL;DR
This systematic review provides empirical evidence suggesting that contemporary AI systems in healthcare introduce privacy and security risks that may challenge traditional assumptions about data protection, and underscores the need for privacy- and security-by-design approaches and governance frameworks that address risks across the AI lifecycle.
Abstract
Background
Artificial intelligence (AI) is increasingly integrated into healthcare information systems, supporting clinical decision-making, imaging analysis, and predictive modeling. While these applications offer operational and clinical benefits, they also introduce emerging risks to patient privacy, data security, and system reliability.
Objective
To systematically review empirical evidence on privacy breaches, security vulnerabilities, and misuse associated with AI applications in healthcare settings.
Methods
PubMed, Embase, Web of Science, Scopus, IEEE Xplore, and ACM Digital Library were searched for empirical studies published between January 2015 and November 2025 that evaluated AI use or misuse in clinical diagnosis, treatment, or decision-making. Two reviewers independently screened studies and extracted data using a standardized form. Findings were synthesized narratively due to heterogeneity in study designs, AI methods, and reported outcomes.
Results
Of 7,285 records identified through database searches and 205 through citation screening, 22 empirical studies met the inclusion criteria, spanning multiple clinical domains and data modalities, predominantly medical imaging applications. Five recurring threat categories were identified: patient re-identification, membership inference, unauthorized access and adversarial exploitation, input manipulation, and misuse or overinterpretation of AI outputs. Across studies, AI models were shown to encode latent biometric signals across diverse data types, limiting the effectiveness of traditional anonymization and synthetic data approaches. Adversarial attacks and input manipulation were also shown to compromise diagnostic performance and system integrity.
Conclusion
This systematic review provides empirical evidence suggesting that contemporary AI systems in healthcare introduce privacy and security risks that may challenge traditional assumptions about data protection. These findings underscore the need for privacy- and security-by-design approaches and governance frameworks that address risks across the AI lifecycle.
Artificial intelligence is moving from experimental use toward routine healthcare infrastructure, yet evidence remains fragmented across clinical, operational, and security domains. This systematic review synthesizes 35 peer-reviewed studies addressing five interconnected applications: predictive analytics, medical imaging, cybersecurity, resource allocation, and clinical decision support. Searches were conducted across PubMed/MEDLINE, IEEE Xplore-indexed records, publisher databases, and Google Scholar for English-language studies published from 2016 through June 2026. Eligible studies were assessed for clinical relevance, validation quality, transparency, bias, implementation readiness, and patient-safety implications. The evidence shows that predictive models can support earlier risk identification, readmission prevention, antimicrobial-resistance forecasting, and population-health surveillance. Imaging systems frequently achieved specialist-level discrimination in constrained datasets, particularly for retinal, dermatologic, breast, lung, and brain imaging, although external validation and workflow integration remained uneven. Cybersecurity research demonstrated value in anomaly detection, federated learning, privacy-preserving analytics, and adversarial-risk management, while also revealing that medical AI creates new attack surfaces. Resource-allocation studies supported demand forecasting, patient-flow management, supply-chain resilience, and treatment optimization. Clinical decision-support systems offered the strongest practical benefit when recommendations were explainable, calibrated, integrated into existing workflows, and subject to clinician oversight. Across all domains, performance declined when models encountered population shifts, incomplete data, weak governance, or poorly designed interfaces. The review concludes that healthcare AI should be evaluated as a sociotechnical intervention rather than a standalone algorithm. Sustainable value depends on prospective validation, equity auditing, cybersecurity-by-design, economic evaluation, transparent reporting, and continuous post-deployment monitoring. These requirements provide a unified roadmap for safer, more effective, and accountable adoption.
Anastasia Andrevna, Druv Juel· International Journal of Art...· 0 citations
Background Biomedical artificial intelligence (AI) requires the integration of privacy-enhancing technologies (PETs) to safeguard sensitive clinical, imaging, and genomic data while preserving analytical utility. Objectives This review critically and systematically maps applications of PETs across the biomedical AI lifecycle in accordance with PRISMA-ScR guidelines and evaluates their technical trade-offs, deployment feasibility, and residual risks. Methods We systematically searched PubMed, IEEE Xplore, ACM Digital Library, and Scopus for studies published between 2015 and 2025. Eligible studies addressed differential privacy, federated learning, secure multiparty computation, homomorphic encryption, or hybrid approaches in biomedical AI. Data were charted on PET type, modality, lifecycle stage, utility metrics, privacy parameters, and deployment considerations. A critical appraisal rubric assessed threat-model adequacy, methodological clarity, reproducibility, privacy–utility transparency, and deployment realism. Additionally, we hand-searched major venues (USENIX Security, NeurIPS, AAAI) and screened Google Scholar for grey literature, applying de-duplication across sources. Results We identified 87 studies spanning clinical decision support, genomics, and medical imaging. From 25,761 initial records, 3,754 underwent title/abstract screening and 1,968 underwent full-text assessment. PETs demonstrated distinct strengths and limitations: differential privacy provided provable guarantees but reduced performance on imbalanced data; federated learning improved data access but remained vulnerable to gradient leakage; and cryptographic methods ensured confidentiality at high computational cost. Synthetic data generation supported privacy-conscious data sharing and benchmarking but remained sensitive to disclosure risk, fidelity loss, and subgroup representation. Hybrid and emerging approaches, including trusted execution environments, zero-knowledge proofs, and privacy-preserving transformer architectures, mitigated composability gaps yet lacked full end-to-end assurance. Case studies at hospital and biobank scale illustrated practical feasibility and infrastructure demands. Conclusions Situating PETs within technical and operational contexts clarifies their capabilities, limitations, and deployment challenges. Residual risks persist, including fairness concerns, inference-time leakage, and overreliance on PETs as compliance proxies. Sustained technical innovation and institutional governance remain essential for the trustworthy integration of PETs in biomedical AI.
Seha Ay, Umit Topaloglu, Wei Zhang· Frontiers in Digital Health· 0 citations
PURPOSE
Artificial intelligence (AI)-enabled software as a medical device (SaMD) is increasingly used across clinical specialties, but its governance remains difficult because adaptive systems raise ongoing concerns about version control, subgroup performance reporting, and postmarket performance drift. This scoping review examined the ethical and regulatory issues surrounding AI-enabled SaMD across jurisdictions and across the product lifecycle, drawing together peer-reviewed literature, standards, and official guidance.
METHODS
A scoping review was conducted using the Joanna Briggs Institute framework and reported in accordance with PRISMA-ScR. Sources published between 2015 and 2025 were identified from PubMed, PubMed Central, Google Scholar, medRxiv/bioRxiv, and the Cochrane Library. Included records were charted by lifecycle stage, jurisdiction, and six prespecified themes: transparency, equity, privacy and security, accountability, lifecycle oversight and change control, and convergence versus fragmentation. Coding was refined iteratively during reviewer calibration.
FINDINGS
Of 21,672 records identified, 369 met the inclusion criteria. Most sources were published from 2019 onward and were concentrated in United States and European Union regulatory settings, with additional contributions from international bodies such as International Medical Device Regulators Forum, WHO, and ISO, as well as from emerging economies including China and India. Postmarket oversight emerged as the most strongly emphasized lifecycle stage, especially in relation to real-world monitoring, drift management, and prespecified change control. Common gaps included limited subgroup reporting, inconsistent expectations for drift thresholds and rollback criteria, and poor alignment between horizontal AI rules and device-specific regulatory frameworks.
IMPLICATIONS
The evidence base shows meaningful progress in the governance of AI-enabled SaMD, but implementation remains uneven across jurisdictions and lifecycle stages. Priority areas include standardized equity reporting, clearer minimum expectations for drift management, and more explicit integration between horizontal AI governance frameworks and SaMD-specific regulatory requirements. These findings support stronger accountability through improved reporting standards, clearer postmarket controls, and better alignment of quality-management processes across the AI-SaMD lifecycle.
Shaharyar Ahsan, Vivian Annastasia Chinyem Obi, D. Noor et al.· Clinical Therapeutics· 0 citations
Current evaluations emphasize accuracy while underreporting privacy, security, robustness, explainability, and verifiability, underscoring the need for comprehensive, guideline-based, multidomain evaluation before deployment in high-stakes clinical settings.
Hikaru Matsuoka, Takayuki Takahashi, Takayuki Semitsu et al.· Online Journal of Public Hea...· 0 citations
Introduction Artificial intelligence (AI) is reshaping healthcare, enabled by advances in computing, affordable data storage, and the widespread adoption of electronic health records (EHRs). Machine learning (ML), deep learning (DL), and natural language processing (NLP) are increasingly used for disease diagnosis, risk prediction, and treatment planning. Objective This systematic review aimed to examine AI applications across clinical domains from 2020 to 2025, assess their diagnostic accuracy and clinical performance relative to standard practice, identify key implementation barriers including regulatory compliance, algorithmic fairness, and transparency challenges, and compare validation practices and methodological quality with earlier systematic reviews. Methods This systematic review followed PRISMA 2020 guidelines. We searched five databases (PubMed, IEEE Xplore, Web of Science, Springer, and Semantic Scholar) for studies published from January 2020 to September 2025. We included original clinical AI studies that reported prospective validation and/or external validation. Results Twenty studies met the inclusion criteria. Publication volume peaked in 2024 (n = 7, 35.0%). DL approaches were most common (n = 12, 60.0%), with convolutional neural networks (CNNs) frequently applied to medical imaging tasks. By clinical domain, 30.0% of studies focused on radiology (n = 6), 20.0% on oncology (n = 4), and 15.0% on cardiology (n = 3). For imaging-based diagnostic models, the descriptive median performance across individual studies was 0.91 AUC (no formal meta-analysis was conducted due to heterogeneity in study designs, populations, and outcome metrics). The most frequently reported challenges were regulatory compliance (55.0%, n = 11), limited algorithmic transparency (40.0%, n = 8), data quality limitations (35.0%, n = 7), and barriers to clinical integration (30.0%, n = 6). Conclusions AI demonstrates strong potential to improve the effectiveness, safety, and quality of healthcare. However, broader clinical adoption remains constrained by regulatory requirements, interpretability gaps, data quality issues, and workflow integration challenges, underscoring the need for stronger validation practices and more implementation-focused research.
Ghulam Hussain Noori, Shaista Bibi, Seung Won Lee· Inquiry : a journal of medic...· 0 citations
Background: AI/ML-enabled medical devices are increasingly deployed in healthcare under evolving regulatory frameworks. As these systems become more integrated into clinical decision-making, there is growing expectation that they demonstrate key dimensions of trustworthy AI to support clinician, patient, and public trust. Whether publicly available regulatory documentation provides sufficient evidence to independently assess the trustworthiness of cleared AI systems remains unclear. Methods: We analysed FDA AI/ML-enabled medical device summary reports published between 2021 and 2025. Reports underwent automated keyword screening followed by multi-stage manual consensus review to identify documented evidence for the six FUTURE-AI principles: Fairness, Universality, Traceability, Usability, Robustness, and Explainability. Descriptive, temporal, and clinical-domain analyses were performed. Multivariable logistic regression assessed whether year of clearance or clinical domain predicted higher reporting transparency, defined as evidence reported for three or more principles. Results: Of 1,105 FDA summary reports screened, 519 were included. Trustworthy AI reporting was limited and uneven. Nearly one quarter (24.7%) provided no evidence for any principle, and none documented evidence across all six. Robustness was most frequently reported (57.6%), while Traceability (8.3%) and Explainability (3.5%) were the most pronounced gaps. Neither year of clearance (OR 1.02, 95% CI 0.88-1.19) nor clinical domain (OR 0.73, 95% CI 0.46-1.15) predicted higher reporting transparency. Interpretation: Substantial, persistent trustworthy AI reporting gaps exist in FDA documentation. Regulatory approval alone should not be considered a proxy for trustworthiness. Standardised, audit-ready reporting across the AI lifecycle is needed to support independent assessment and responsible adoption of healthcare AI.
Ahmed M. A. Salih, Oliver Díaz, Alejandro Guzmán et al.· 0 citations