Skip to content
Preprint

GRASP: Reinforcing Language Model Anonymizers with Group Relative Policy Optimization

Aug 2026 · 0 citations · 28 references
Computer Science

TL;DR

A single small model acts as anonymizer, adversary, and utility judge, trained against a self-generated reward that hides attributes while preserving meaning, with a design that guards against reward hacking.

Abstract

Large language models can infer sensitive personal attributes, such as age, location, and occupation, from ordinary text, turning everyday writing into a privacy risk. Adversarial anonymization defends against this by rewriting a text with a capable language model that also plays the attacker, but it needs a powerful model at inference time and thus sends private text to a third party, the very exposure anonymization should prevent. Recent work distills this behavior into a small on-device model using supervised fine-tuning and direct preference optimization (DPO), but DPO only imitates the teacher's offline choices and never directly optimizes the privacy--utility objective we care about. We introduce \textbf{GRASP} (\textbf{G}roup-\textbf{R}elative \textbf{A}nonymization via \textbf{S}elf-refinement \textbf{P}olicy-optimization), which reinforces the local anonymizer online with Group Relative Policy Optimization. A single small model acts as anonymizer, adversary, and utility judge, trained against a self-generated reward that hides attributes while preserving meaning, with a design that guards against reward hacking. Trained on Llama-3.1-8B, \ours{} improves the privacy--utility trade-off over the DPO-distilled baseline, consistently across three independent LLM judges. Against adversarial anonymization driven by frontier models such as Gemini~2.5~Flash and Claude, it achieves a comparable or better overall trade-off while removing substantially more private information, and it runs entirely on-device at roughly $1\%$ of the GPT-4o teacher's cost.

View source

Similar papers

Preprint Aug 2026

Personalized Privacy Control in LLMs via Attention Head Intervention

The introduction of personalized privacy, which incorporates user-specific disclosure preferences into privacy control, and a robust inference-time attention head intervention method that adjusts disclosure behavior toward policy-consistent responses are proposed.

Junseok Kim, Nakyeong Yang, Kyomin Jung · 0 citations
Preprint Aug 2026

Privacy Without Regret: Differentially Private Inference-Time Alignment

Private Inference-Time Pessimism (PrivITP) is introduced, which combines $\chi^2$-regularized rejection sampling with a two-phase Gaussian mechanism, and achieves ex-post $(\epsilon,\delta)$-DP with a privacy cost independent of the number of responses, cleanly decouples the regularization parameter from the privacy pa...

I. Jain, Nandini Bhattad, Sayak Ray Chowdhury · 0 citations
2026

PI-SAFE: Practical Privacy-Preserving LLM Inference With Adversarial Fine-Tuning for Optimized Utility

Cloud-based Large Language Model (LLM) inference services typically require users to submit plain-text inputs, thereby posing severe privacy risks. Existing privacy-preserving paradigms are mostly task-specific and often necessitate pervasive modifications to the entire server-side model. This reliance introduces subst...

Wentao Zhong, Yu-Ting Li, Di-Cong Yu et al. · 0 citations
#machine learning Review Sep 2026

Differentially Private Semantic Plans for Aggregate Insight Generation

\texttt{URANIA} provides end-to-end differential privacy (DP) for summaries of data-dependent clusters. However, its cluster--keyword release does not directly provide collection-wide aggregates for semantic concepts defined independently of the protected corpus. Records may express several concepts, records expressing...

Behrooz Razeghi · 0 citations
#artificial intelligence Preprint Sep 2026

EP-Mem: Elastic Privacy Memory for Social Relationship-Aware LLM Agents

Large language model (LLM) agents face critical privacy risks when acting as delegates in human-agent-human communication. To prevent such breaches, agents must understand users'social relationships and adhere to context-dependent social information disclosure boundaries. Current studies on agent memory privacy focus o...

Feng-Zhou Sun, Yuan Zhang, Xin-Tong Yu et al. · 0 citations
Preprint Sep 2026

Inferring Hidden User Models from the Behavior of Personalized LLM Agents

Recent personalized LLM agents increasingly transform information retained in memory into compressed or structured representations, which we call user models, to guide later decisions. When source wording is removed from the state reachable through the ordinary interface, these models are commonly treated as more priva...

Hao-Yang Li, Ya-Xin Xiao, Qing-Qing Ye et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.