A novel metric, ``Dependency Triad''(DT), is proposed, which summarizes the pairwise dependency information relevant to CPL using three parameters and yields a conservative estimator of pairwise CPL, which is particularly suitable for high-cardinality attributes.
Abstract
Collecting multidimensional user data is essential for extracting rich insights across various applications. Local Differential Privacy (LDP) has emerged as a de facto standard for mitigating privacy risks in such scenarios. A key challenge in privacy-preserving multidimensional data collection lies in inter-attribute dependencies, as they can inadvertently reveal correlated information and increase privacy vulnerabilities. Therefore, accurately measuring correlation-induced privacy leakage (CPL) is essential for privacy analysis and privacy-utility trade-off. However, existing CPL analysis solutions either require accurate prior knowledge or face scalability challenges for large numbers of attributes and high-cardinality attributes. These limit their practical applicability in real data. To address this research gap, we propose a novel metric, ``Dependency Triad''(DT), which summarizes the pairwise dependency information relevant to CPL using three parameters and yields a \emph{constant-time} conservative estimator of pairwise CPL. DT explicitly models uncertainty in prior distributional knowledge through its parameters, delivering robust leakage estimates. Moreover, its robustness to sparse distributions makes it particularly suitable for high-cardinality attributes, while the pairwise formulation serves as a tractable building block for assessing total leakage in multidimensional settings. Extensive experiments on both synthetic and real datasets demonstrate that DT consistently estimates CPL across diverse dependency regimes and prior uncertainties.
CoP is proposed, a coordinated perturbation mechanism designed to mitigate CIL in multidimensional data collection while preserving utility and significantly outperforms state-of-the-art LDP mechanisms in reducing disclosure while preserving analytical accuracy.
Sandaru Jayawardana, Ming Ding, Kanchana Thilakarathna· Proceedings on Privacy Enhan...· 0 citations
: As data spaces emerge to facilitate sovereign data exchange, geospatial data has become a critical resource across various domains. However, sharing sensitive geodata, such as cadastral property records, presents a privacy-utility trade-off. While the European Data Governance Act (DGA) establishes data trustees as the technical intermediaries authorized to perform necessary anonymization, a gap remains between theoretical spatial privacy algorithms and their practical, policy-driven application. This research addresses this gap by presenting an empirical evaluation of 11 spatial anonymization methods applied to complex polygon geometries. Using a dataset of 2,147 forested cadastral parcels, we quantify the trade-offs using privacy metrics ( k -anonymity, differential privacy ε ) and utility measures (Hausdorff distance, area deviation, centroid shift). Our results identify algorithmic failure modes, such as the “sparse forest” phenomenon, and demonstrate the improved performance of combined hybrid approaches. To operationalize these findings, we integrate our results into the architecture of data trustees operating within various domain-specific data spaces. We demonstrate how the empirically derived, use-case-specific anonymization guidelines can be translated into machine-interpretable Open Digital Rights Language (ODRL) policies. By mapping spatial transformation algorithms to ODRL constraints, this research provides data trustees with an approach to automatically enforce spatial privacy and sovereignty within data spaces.
Michael Steinert, Bekzod Nazarov, Thorsten Reitz et al.· Proceedings of the 15th Inte...· 0 citations
ProxyDrift is presented, a framework that identifies and measures drift between production traffic and offline evaluation sets, and constructs and refreshes those evaluation sets accordingly; all without access to raw user data.
Michael Levit, Josh Ledgard, Haoyu Dong et al.· 0 citations
This work investigates a probabilistic variant of PCD, where an LLM-driven probabilistic estimation of k-anonymity is augmented with an LLM-driven probabilistic estimation of k-anonymity, and proposes k-anonymity as a useful auxiliary metric for tackling PCD.
With the widespread adoption of graph-structured data, protecting the complex relational information between nodes and edges while preventing sensitive information leakage has become a critical challenge. However, existing edge protection methods either introduce noise directly into the adjacency matrix, resulting in significant information loss, or uniformly apply noise across all edges, leading to imbalanced privacy budget allocation and inefficiency. To address these issues, we propose DPEI, a Differential Privacy-based Edge Information protection solution designed to safeguard the edge relationships between two nodes, thus reducing the risk of privacy leakage and preventing attackers from repeatedly inferring internal community relationships from the released graph data. Specifically, DPEI achieves protection through PPO (Proximal Policy Optimization)based selection of locally optimal thresholds combined with adaptive Laplace noise operations, and attachment nodes below the threshold into high-information edges to enhance relational information protection. Subsequently, unlike traditional uniform allocation, DPEI distributes the privacy budget in proportion to the information content of each edge, ensuring that edges with higher information content receive stronger privacy protection. Extensive experiments conducted on three real-world graph datasets demonstrate that DPEI significantly outperforms existing methods across seven commonly used graph metrics, thereby validating its effectiveness and practicality.
Privacy-preserving data publishing and explainable artificial intelligence (XAI) are both essential for trustworthy machine learning, yet their interaction remains largely underexplored. In practice, models are often trained on anonymized datasets, but little is known about how classical anonymization techniques affect post-hoc explanations. In this paper, we provide a systematic empirical study of how feature attribution rankings change under widely used anonymization models, including k-anonymity, $\ell$-diversity, t closeness, and $(\alpha, k)$-anonymity. Across multiple real-world datasets and classifiers, we compare explanations generated by SHAP and LIME and quantify their stability using rank correlation and hypothesis testing. Our findings reveal a fundamental trade-off: explainable privacy-preserving models are feasible under mild privacy constraints, but strict anonymization requirements often lead to unstable explanations and severe utility degradation.
Casper Lauge Nørup Koch, Mina Alishahi, Gaurav Choudhary· 2026 IEEE European Symposium...· 1 citation