Skip to content
Preprint

P2Skill: Privacy Preserving Skill Distillation for Cloud-Local LLM Inference Systems

Aug 2026 · 1 citation · ⚡ 1 influential · 44 references
Computer Science

TL;DR

P2Skill is proposed, a prompt-based skill distillation method in which a local small language model (SLM) autonomously performs decomposition, PII-aware routing, paraphrasing, and reconstruction by following the skill prompts.

Abstract

Cloud-local LLM inference systems have the potential to use the reasoning capability of large cloud models while protecting sensitive user data on personal devices. Cloud-bound requests must exclude personally identifiable information (PII) to prevent external data leakage. Existing privacy-preserving methods rely on prompt perturbation, entity masking, or model fine-tuning, but these approaches may distort contextual semantics or require additional training. This paper proposes P2Skill, a prompt-based skill distillation method in which a local small language model (SLM) autonomously performs decomposition, PII-aware routing, paraphrasing, and reconstruction by following the skill prompts. Skills are iteratively refined from execution failures by a cloud LLM, enabling the local SLM to generalize beyond memorized PII patterns, and therefore P2Skill requires no privacy-specific fine-tuning or learned auxiliary detectors. Evaluation on a four-domain benchmark shows that P2Skill achieves $1.69\times$ and $3.66\times$ higher privacy-preserved inference quality than previous baselines.

View source

Similar papers

2026

PI-SAFE: Practical Privacy-Preserving LLM Inference With Adversarial Fine-Tuning for Optimized Utility

Cloud-based Large Language Model (LLM) inference services typically require users to submit plain-text inputs, thereby posing severe privacy risks. Existing privacy-preserving paradigms are mostly task-specific and often necessitate pervasive modifications to the entire server-side model. This reliance introduces subst...

Wentao Zhong, Yu-Ting Li, Di-Cong Yu et al. · 0 citations
#machine learning Preprint Sep 2026

Privacy-Preserving Split Learning for Federated LLM Fine-Tuning

This work addresses leakage through a learned obfuscate-and-recover scheme that protects participants' private datasets while still allowing an independently deployable model to be trained on the server side, making split-based federated LLM fine-tuning practically viable.

Heng Jin, Chao-Yu Zhang, He-Xuan Yu et al. · 1 citation
#artificial intelligence Preprint Aug 2026

Decoupling Knowledge and Privacy: Post-Task Self-Distillation Replay for LLM Continual Learning

Privacy-preserving continual learning (PPCL) must reduce the reproduction of sensitive content while retaining useful knowledge across sequential tasks. Formal privacy guarantees characterize randomized mechanisms, whereas operational output control concerns whether a trained model selectively reduces the likelihood of...

Sheng-Tao Wen, Yun-Ying Yang, Xiang Chen et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Learning to Refer: Client-Resolved Generation for Privacy-Aware Language Models

Cloud-based large language models (LLMs) require users to disclose plaintext data to service providers, creating privacy risks in sensitive domains. Existing privacy-preserving approaches often trade utility for protection, incur substantial computational or communication overhead, remain vulnerable to reconstruction f...

Jeongho Yoon, Chanhee Park, Yong-Chan Chun et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Auditing and Mitigating Privacy Leakage in Cloud-Edge Collaborative Decoding

CoVeil is proposed, a defense mechanism which dynamically optimizes transmitted signals to suppress leakage during decoding time while preserving the collaborative quality, and consistently improves the privacy-utility trade-off over existing baselines by reducing data leakage.

Ke-Jia Zhang, Tianyuan Zou, Zi-Xuan Gu et al. · 0 citations
Preprint Aug 2026

Gecko: Fast Private Inference via Secure Public Encoder Offloading

Gecko is presented, designed to limit this additional risk while retaining a compact encrypted predictor, and formalizes ideal independence and information-preservation conditions as design guidance, then separately evaluate component-reuse extraction attacks.

Cheng'an Wei, Kai Chen, Yue Zhao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.