A New Multidimensional Data Anonymization Algorithm for Privacy-Preserving Data Publishing
Developing a privacy-preserving data publishing algorithm that prevents individuals’ identity information from being disclosed without ignoring the utility of the data is still an important goal to be achieved. Finding the optimal balance between data utility and data privacy is an NP-hard problem. In this study, a new k-anonymization-based data anonymization algorithm is proposed. The proposed algorithm partitions the data space with a new efficient partitioning strategy based on the k-dimensional tree (KD-tree), resolves the boundary problem while maintaining the trade-off balance, and introduces a new mechanism to address the problems caused by outliers. Moreover, it can be applied to both numerical and categorical data. The experimental results indicate that the proposed algorithm achieves competitive or improved performance compared with the baseline algorithms across seven commonly used evaluation metrics. Overall, the findings suggest that the proposed framework can improve the utility–privacy trade-off in multidimensional k-anonymization.