Skip to content

Privacy-Constrained Distributionally Robust Detection and Collaborative Attribution of LLM-Generated Text

Sep 2026 · International journal of pattern recognition and artificial intelligence · 0 citations
Authorship Attribution and Profiling Topic Modeling

TL;DR

FedRAT is proposed, a privacy-preserving federated framework that couples paired paraphrase consistency, generator-family auxiliary supervision, differentially private client updates, and loss-dependent aggregation within a unified detection setting.

Abstract

With the widespread use of large language models, distinguishing human-written text from LLM-generated text has become increasingly important for content authenticity, academic integrity, and digital forensics. However, existing detectors remain vulnerable to paraphrase attacks, domain shift, and privacy restrictions that prevent institutions from sharing raw text data. To address these challenges, this paper proposes FedRAT, a privacy-preserving federated framework that couples paired paraphrase consistency, generator-family auxiliary supervision, differentially private client updates, and loss-dependent aggregation within a unified detection setting. FedRAT jointly learns binary AI-text detection and weak attribution to a predefined generator family through a shared encoder. Experiments on multi-domain and multi-generator benchmarks show that FedRAT consistently outperforms local training, standard federated learning, paraphrase-augmented federated learning, and representative detection baselines under clean and paraphrased settings. The ablation results evaluate attribution learning, paraphrase consistency, risk-aware aggregation, and differential privacy, while Expected Calibration Error and Brier Score provide diagnostic measures of confidence quality. The results support FedRAT as an empirically evaluated integration for privacy-constrained LLM-generated text detection and coarse generator-family attribution under the tested rewriting conditions.

View source

Similar papers

2026

PI-SAFE: Practical Privacy-Preserving LLM Inference With Adversarial Fine-Tuning for Optimized Utility

Cloud-based Large Language Model (LLM) inference services typically require users to submit plain-text inputs, thereby posing severe privacy risks. Existing privacy-preserving paradigms are mostly task-specific and often necessitate pervasive modifications to the entire server-side model. This reliance introduces subst...

Wentao Zhong, Yu-Ting Li, Di-Cong Yu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Learning to Refer: Client-Resolved Generation for Privacy-Aware Language Models

Client-Resolved Generation (CRG), a genera- tion interface that separates server-side generation from the lexical realization of input-derived content, is introduced, which provides a practical interface for privacy-sensitive cloud LLMs by reducing plaintext exposure across both input and output pathways while preserv-...

Jeongho Yoon, Chanhee Park, Yong-Chan Chun et al. · 0 citations
#natural language process... Preprint Sep 2026

Conformal Privacy Auditing: Calibrated Re-identification Attacks with Statistical Guarantees

Conformal Privacy Auditing is introduced, a distribution-free calibration framework that provides a statistical certificate of re-identification risk for each released document against LLM-empowered adversaries and enables audits of open-source models and proprietary API models in a unified framework.

Shuo Huang, G. Haffari, Xing-Liang Yuan et al. · 0 citations
#machine learning Preprint Sep 2026

QuanText: Protecting Dataset-Level Secrets in Textual Data Sharing

Natural-language datasets support many downstream applications and research studies, but releasing text can reveal sensitive global properties of the underlying data source, such as the proportion of records associated with a particular gender, diagnosis, or political stance. Existing work has largely focused on proper...

Shuai-Qi Wang, Zinan Lin, Giulia Fanti · 0 citations
Preprint Aug 2026

Beyond Visual Evidence: Revealing and Mitigating Relational Privacy Leakage in Document MLLMs

The Dynamic Relational Unlearning Framework (DRUF) is proposed which comprises a Relational Decoupling Unlearning (RDU) module and a dynamic set update mechanism that suppresses the leakage of high-risk field pairs while preserving KIE performance.

Bei-Ning Xu, Hairui Wang, Jiaxin Wang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Auditing and Mitigating Privacy Leakage in Cloud-Edge Collaborative Decoding

CoVeil is proposed, a defense mechanism which dynamically optimizes transmitted signals to suppress leakage during decoding time while preserving the collaborative quality, and consistently improves the privacy-utility trade-off over existing baselines by reducing data leakage.

Ke-Jia Zhang, Tianyuan Zou, Zi-Xuan Gu et al. · 0 citations

Related blog posts

GPT-Lab Aug 28, 2026

We built an AI factory for HVAC control

What does it take to trust AI-driven HVAC optimization? Our AI Model Factory combines agents, machine learning, reinforcement learning and deterministic checks in a governed workflow designed for messy, real-world building data. The post We built an AI factory for HVAC control appeared first on GPT-Lab.

Microsoft Research Blog Jul 30, 2026

EvoLib: Turning experience into evolving knowledge

LLMs do not get smarter just by remembering more. EvoLib turns experience into evolving knowledge, taking reusable skills and insights that help models learn and adapt across tasks long after deployment. The post EvoLib: Turning experience into evolving knowledge appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.