Skip to content
Conference

Quantifying and Preventing AI-Aware Privacy Leakage in Large-Scale Data Ecosystems: A Systematic Review

Jul 2026 · 2026 7th International Conference on Smart Systems and Inventive Technology (ICSSIT) · pp. 1321-1325 · 0 citations · 27 references

Abstract

Now, AI runs on cloud platforms, edge systems with federated settings, and in large language model (LLM) pipelines or data-sharing services, creating even wider privacy leakage paths beyond classical database disclosure. This paper offers a systematic, structured review of the literature on a curated, cost-effective reference corpus for quantifying and preventing privacy leakage in AI-enabled data ecosystems. The review ties together four strands of research that are often treated separately. Firstly, the privacy risk throughout the AI life cycle. Secondly, the measurement of the quantitative leakage. Thirdly, architectures of the privacy-preserving models, and finally, operational governance for real-world deployment. Our analysis demonstrates that state-of-the-art approaches are moving from static mechanisms based on anonymization to metric-aware protections, including information-theoretic leakage scores, cumulative differential privacy accounting, personalized privacy budgets, and benchmark-driven attack evaluation. In parallel, prevention methods are evolving beyond single homomorphic noise injection and are becoming multi-layered defenses that combine differential privacy, federated learning, weight quantization, synthetic data generation, policy-driven automation, and LLM controls. The review uncovers four itchy gaps: fractured assessment metrics, shaky privacy-utility trade-offs, flimsy integration of technological controls and compliance processes, and low cross-context validation across cloud-based computing, edge computing (data processing at or near the source), federated learning (distributed machine-learning methods), and generative AI systems. The paper concludes by outlining a unified research agenda to build AI-aware, quantifiable, and usable privacy protection stacks.

View source

Similar papers

Conference Jul 2026

Privacy Protection in Cross-System Data Exchange: A Comprehensive Review of Multi-Party Computation Approaches

In today’s era of big data, personal privacy is increasingly at risk due to widespread data sharing. Mobile applications often collect excessive personal information, while advanced analytics can sometimes lead to biased or discriminatory practices. These challenges create an urgent need for secure, privacy-preserving methods that allow sensitive data to be shared and analyzed across multiple parties and diverse systems. This paper reviews the progress made in this area, with a particular focus on the requirements for safe data sharing and controlled dissemination of private information during multi-party data fusion. The review is structured around three main perspectives: privacy-preserving computation, information sharing control, and collaborative secure computation. We begin by examining the current state of privacy protection in large-scale, interconnected environments, followed by a comparison of recent research developments at both national and international levels. In the area of privacy-preserving computation, emerging techniques such as full-lifecycle privacy safeguards, information flow control, and secure data exchange mechanisms are discussed. For information sharing control, three approaches are analyzed—local control, extended control, and desensitization methods. In collaborative secure computation, we outline methods currently being applied in both academic and industry contexts. Finally, the paper highlights key challenges and directions for future research. Traditional approaches such as anonymization, perturbation, and access control, as well as more advanced methods like cryptography and federated learning, all face practical limitations. To achieve robust protection throughout the entire data lifecycle, theoretical models and privacy-aware information systems must be further refined and tailored to different real-world application scenarios.

Chaitanya Tumma, Supraja Ayyamgari, Charan Thumma et al. · 0 citations
Conference Open access 2026

A Serverless Client-Side Privacy Index for Sensitive Data Processing

: The widespread adoption of data-driven systems has intensified concerns regarding the protection of sensitive information and the assessment of privacy risks. Although numerous privacy-preserving techniques and models have been proposed, quantifying and interpreting the level of privacy achieved remains challenging, particularly for non-expert users. This paper introduces the Privacy Index, a unified metric designed to aggregate multiple privacy-related factors into a single, interpretable score. The proposed approach integrates the validation of established privacy models, detection of anonymization techniques, and estimation of re-identification risk under different attacker assumptions. To support practical usage, we develop a serverless, client-side web application that automatically processes structured datasets by classifying attributes and computing the Privacy Index without transmitting data externally. Experimental evaluation demonstrates that the attribute classification component achieves average confidence levels above 88% across multiple datasets, correctly identifying all direct identifiers with high reliability. The system successfully validates privacy models such as k -anonymity and l -diversity and effectively distinguishes between poorly and well-anonymized datasets. The tool is open-source, its code is available at https://github.com/ieeta-mith/DataPrivScore and a demo is available through https://ieeta-mith.github.io/DataPrivScore/.

José A. Gameiro, J. Oliveira, João Rafael Almeida · 0 citations
#diffusion models Open access Oct 2026

Measuring Legislature-Aligned Privacy Risks in Synthetic Graphs

SyntheGrAnon is introduced, a framework for evaluating synthetic graph anonymity that primarily targets the singling out, linkability, and inference risks outlined in the EU GDPR at the node and community levels, while also including edge-level attacks as an extension of the node-level setting.

Abele Malan, Ahmad Al Kurdi, Stefanie Roos et al. · 0 citations
Open access 2022

Federated Analytics for Privacy-Preserving Edge Computing

With the rapid expansion of edge computing, vast volumes of sensitive data are now being generated and processed at the network's periphery, raising significant concerns about privacy and data security. Federated Analytics (FA) emerges as a transformative solution by enabling decentralized data analysis without the need to transfer raw data to central servers, thereby mitigating potential privacy breaches. This study investigates the integration of FA into edge computing ecosystems, leveraging advanced Privacy-Enhancing Technologies (PETs) such as Differential Privacy (DP), Secure Multiparty Computation (SMC), and Homomorphic Encryption (HE) to ensure robust privacy protections. A multi-layered architecture is proposed and evaluated using simulations on Raspberry Pi clusters and synthetic workload datasets to emulate real-world edge environments. Experimental results indicate that FA, especially when combined with DP, achieves a strong balance between analytical accuracy and computational efficiency, while SMC and HE offer enhanced security at the cost of increased computational overhead. The findings underscore the practicality and effectiveness of FA for privacy-preserving analytics at the edge, suggesting its potential to support compliance with data protection regulations and meet the demands of future applications. The paper concludes by emphasizing the need for further research in optimizing scalability, minimizing resource usage, and exploring synergies with emerging technologies such as 6G and intelligent orchestration platforms to fully realize the promise of federated edge analytics.

John McCarthy, M. Minsky · 0 citations
Conference Aug 2026

Big data privacy protection and data security based on deep learning

A federated deep learning framework that systematically integrates adaptive privacy noise mechanisms and trust-weighted aggregation within a distributed architecture that ensures the protection of sensitive data during collaborative analysis through precise differential privacy control and advanced neural network models is presented.

Changwen Xu, Yanqing Ding, Rui Ding · 0 citations
Open access 2018

A Framework for Privacy-Preserving Machine Learning in Sovereign Cloud Environments

The rapid growth of machine learning (ML) technologies has raised concerns about the privacy and security of sensitive data used in training models. Privacy-preserving techniques such as federated learning, homomorphic encryption, and differential privacy are emerging solutions to protect data in ML applications. However, these techniques often face challenges in terms of scalability, performance, and compliance with data privacy regulations. Sovereign Cloud environments, characterized by strict data governance and jurisdictional controls, offer a potential solution for addressing these challenges. This paper presents a novel framework for integrating privacy-preserving ML techniques within Sovereign Cloud infrastructures. By combining cutting-edge cryptographic approaches with the data sovereignty features of Sovereign Clouds, our framework ensures data privacy, legal compliance, and efficient machine learning at scale. We discuss key challenges in data privacy, scalability, and legal compliance, and propose a set of best practices for deploying privacy-preserving ML in these environments. Additionally, we evaluate the proposed framework through case studies, demonstrating its potential in sectors such as healthcare and finance. The results show that our framework provides a balanced approach to privacy, scalability, and performance, contributing to the future of secure and responsible ML deployment.

Ahmed Hassan · 0 citations