Aug 2026· WIREs Data Mining and Knowledge Discovery· 0 citations· 57 references
TL;DR
This survey provides the first comprehensive and systematic review of text anonymization methods published between 2020 and 2025, covering 48 primary studies identified through a structured search and rigorous screening procedure and reveals a growing shift from identifier‐centric de‐identification toward context‐aware anonymization.
Abstract
Text anonymization has become a critical requirement across healthcare, legal, financial, and online communication domains, where large volumes of sensitive textual data are increasingly used for analytics, information retrieval, and model training. Despite decades of research, anonymization remains challenging due to the diversity of sensitive information types, domain‐specific annotation schemes, and the emergence of complex inference risks that extend beyond explicit identifiers. On the other hand, we have seen a large number of approaches using transformers in recent years. This survey provides the first comprehensive and systematic review of text anonymization methods published between 2020 and 2025, covering 48 primary studies identified through a structured search and rigorous screening procedure. We analyze anonymization approaches across rule‐based, statistical, neural, transformer‐based, and large language model (LLM) paradigms, and compare them across multiple domains, languages, and sensitive‐information taxonomies. Our review also synthesizes available datasets, software resources, and evaluation practices used to assess both privacy protection and text utility. The findings highlight significant fragmentation across domains, persistent limitations in handling quasi‐identifiers and semantic leakage, and substantial inconsistencies in evaluation protocols. We identify key methodological trends, gaps, and emerging challenges, including the integration of LLMs, multilingual settings, and adversarial evaluation. Our analysis reveals a growing shift from identifier‐centric de‐identification toward context‐aware anonymization, while evaluation methodologies have not yet evolved at the same pace. We outline open research directions and propose a roadmap toward more robust, context‐aware, and empirically grounded anonymization systems. This survey aims to establish a unified reference point for researchers and practitioners working on privacy‐preserving NLP and sensitive text processing.
This paper empirically evaluates ChatGPT 3.5 and 4.0 using over 23,000 real user-generated medical queries, assessing their susceptibility to privacy breaches through quasi-identifiers such as age, location, phone number and national registration number and proposes a scalable privacy evaluation model that combines k-anonymity, l-diversity, t-closeness, entropy, re-identification risk and delta-disclosure.
Foad Jalali, Mehran Alidoost Nia· Journal of Supercomputing· 0 citations
This review provides researchers and practitioners with a structured framework for method selection based on their specific constraints and identifies six prioritized research directions for future investigation, identifying critical research gaps including the preservation of multi-word clinical concepts, scarce evaluation in real-world clinical workflows, and persistent hallucination risks in model-based extraction.
This comprehensive review examines the dual role of LLMs in both facilitating and mitigating various information integrity challenges, including misinformation, disinformation, fake news, social bots, and privacy concerns, and demonstrates critical gaps in current approaches.
A systematic evaluation of several proprietary and open-source Large Language Models for sensitive entity extraction from documents and a form-filling pipeline that uses vision-capable LLMs to label form fields, generate realistic synthetic personas, and fill real blank forms, enabling reproducible evaluation across diverse layouts.
Errita Xu, Stefan Larson, Kevin Leach· Proceedings of the 2026 ACM...· 0 citations
Privacy-preserving data publishing and explainable artificial intelligence (XAI) are both essential for trustworthy machine learning, yet their interaction remains largely underexplored. In practice, models are often trained on anonymized datasets, but little is known about how classical anonymization techniques affect post-hoc explanations. In this paper, we provide a systematic empirical study of how feature attribution rankings change under widely used anonymization models, including k-anonymity, $\ell$-diversity, t closeness, and $(\alpha, k)$-anonymity. Across multiple real-world datasets and classifiers, we compare explanations generated by SHAP and LIME and quantify their stability using rank correlation and hypothesis testing. Our findings reveal a fundamental trade-off: explainable privacy-preserving models are feasible under mild privacy constraints, but strict anonymization requirements often lead to unstable explanations and severe utility degradation.
Casper Lauge Nørup Koch, Mina Alishahi, Gaurav Choudhary· 2026 IEEE European Symposium...· 1 citation
This work expands upon the privacy threat assessment model to quantitatively evaluate the risks of data likability, identifiability, non-repudiation, detectability, unintended disclosure, indulgence, and policy & consent noncompliance, and constructs a framework aimed at mitigating these identified risks.
Jamila Alsayed Kassem, Tim Müller, Christopher A. Esterhuyse et al.· 0 citations