Skip to content
Preprint

AEGIS: Attention-Embedding Gradient Isolation Shield - Triple-Channel Gradient Masking for Privacy-Preserving Federated LLM Fine-Tuning

Aug 2026 · 0 citations
Computer Science

TL;DR

AEGIS (Attention-Embedding Gradient Isolation Shield), a lightweight defence that reduces token recovery rates to near zero against a range of gradient inversion attacks, both analytical and optimisation-based, while preserving or improving model utility.

Abstract

Gradient inversion attacks recover private training text from gradients shared in federated learning, posing a serious threat to collaborative model training. Through our analysis of transformer gradient structure, we identify three channels through which private token information leaks: the attention output projection gradient exposes a low-rank subspace that encodes input embeddings (Channel 1), the embedding gradient's row-norm sparsity directly reveals which tokens are present (Channel 2), and the MLP expansion gradient carries a recoverable subspace signal analogous to Channel 1 (Channel 3). State-of-the-art attacks exploit these channels analytically to achieve near-exact token recovery in seconds. Existing defences address at most one channel and either degrade model utility or leave the remaining structural signals intact. We introduce AEGIS (Attention-Embedding Gradient Isolation Shield), a lightweight defence that closes all three analytical channels with three backward-path operations requiring no architectural changes: freezing attention projection parameters eliminates Channel 1 by construction, calibrated noise injection into the embedding gradient destroys Channel 2's token-presence signal, and analogous per-block noise injection into the MLP expansion gradient masks Channel 3. The same masked gradient drives both the local optimiser step and the server export, so no clean signal is retained on either side. Evaluated across 11 models and six datasets, AEGIS reduces token recovery rates to near zero against a range of gradient inversion attacks, both analytical and optimisation-based, while preserving or improving model utility. We provide formal guarantees for Channels 1 and 2 and validate the full defence empirically against adaptive adversaries with complete knowledge of the mechanism.

View source

Similar papers

Preprint Sep 2026

Where to Defend? Layer-Wise Adversarial Training for Robust Transformer-Based Semantic Communications

Deep learning-based semantic communication (DeepSC), a Transformer-based encoder-decoder, achieves semantic fidelity over noisy channels but remains vulnerable to adversarial perturbations injected at multiple stages of the pipeline. We present a layer-wise robustness framework that compares fast gradient sign method (...

Maria Slim, Razane Tajeddine, Mariette Awad et al. · 0 citations
Preprint Sep 2026

MROP: Mask-Region Optimized Purification Against Backdoor Attack in Deep JSCC

Deep joint source and channel coding (JSCC) transmits a source by mapping it directly to channel symbols through an end-to-end deep neural network (DNN) and reconstructing it at the receiver. Taking image transmission as an application, this DNN pipeline behaves as a black box: the receiver cannot readily detect securi...

Seongkyu Yang, Hyeonho Noh, Hyun Jong Yang et al. · 0 citations
Preprint Sep 2026

Your Model Is Leaking: Covert Information Transfer through LLM Residual Streams

Privacy-sensitive organizations may run large language models (LLMs) in restricted or air-gapped environments while exporting selected diagnostic artifacts. We show that a compromised runtime component can hide sensitive information in intermediate activations that are allowed to leave the restricted environment. An of...

Ming-Yuan Li, Yan-Na Jiang, Guang-Sheng Yu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Position Matters: Feature Inversion Attacks in ViT Split Inference with Token Reduction and Shuffling

This work shows that, despite disrupting the spatial structure required by conventional reconstruction attacks, transmitted token embeddings retain substantial positional information, and introduces the Spatially Aligned Reconstruction Attack (SARA), a unified pipeline that predicts token positions, restores their spat...

Stefano Leggio, Giulio Rossolini, Alessandro Biondi · 0 citations
Preprint Aug 2026

SecureDrive-FL: Joint Differential Privacy and Gradient-Aware Selective Homomorphic Encryption for Federated Driver Monitoring

This work introduces GASHE (Gradient-Aware Selective Homomorphic Encryption), a novel selective encryption strategy that dynamically identifies and encrypts only the gradient components exceeding a DP-calibrated sensitivity threshold, rather than encrypting all parameters uniformly as in static layer-based or full-para...

Baran Can Gül, Hanuma Siddhartha Tunuguntla, Anjana Arvind Naik et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.