Skip to content
Preprint

Why Are LLM Backdoor Defenses Fragmented? A Feature-Level Explanation with Sparse Autoencoders

Aug 2026 · 0 citations · 37 references
Computer Science

TL;DR

Results show that SAE-based analysis can explain defense fragmentation and guide interpretable backdoor mitigation and show system?atic encoding differences: dirty-label backdoors are dominated by isolated interaction features, whereas clean-label backdoors rely more on heterogeneous mixtures of mixed and weight-modified features.

Abstract

Backdoor attacks pose a serious threat to large language models (LLMs), but existing defenses remain fragmented, failing to pro?vide unified defense against both dirty-label and clean-label attacks. To investigate why such fragmentation arises, we present the first systematic feature-level mechanistic analysis of LLM backdoors using sparse autoencoders (SAEs). Starting from a 2 x 2 comparison of clean and poisoned models on clean and triggered inputs, we trace backdoor-induced logit shifts to high-contributing SAE features and categorize them into four roles: interac?tion, suppressed, mixed, and weight-modified features. This taxonomy reveals system?atic encoding differences: dirty-label back?doors are dominated by isolated interaction features, whereas clean-label backdoors rely more on heterogeneous mixtures of mixed and weight-modified features. These differ?ences explain why existing defenses remain fragmented across attack paradigms. We val?idate this hypothesis through inference-time feature clamping, which reduces ASR to at most 10.8% in most dirty-label settings and at most 15.4% in the majority of clean-label settings, while preserving benign-task perfor?mance. These results show that SAE-based analysis can explain defense fragmentation and guide interpretable backdoor mitigation.

View source

Similar papers

Preprint Sep 2026

MechAudit-40: White-Box Auditing across 40 LLM Attack Mechanisms

While LLM attacks span prompt optimization, multi-turn context manipulation, retrieval poisoning, and model backdoors, white-box defenses are typically evaluated on isolated attack families. Consequently, whether heterogeneous attacks leave internal representation shifts that generalize to unseen threat mechanisms rema...

Zhen Guo, Shang-Hao Shi, Shamim Yazdani et al. · 0 citations
Preprint Sep 2026

AlcaTRAz - Anchored Tree-Rule Defense Against Jailbreaks

AlcaTRAz (Anchored Tree-Rule defense Against jailbreaks), a prompt-level defense based on rule trees that operates exclusively on the input text and requires no modification or retraining of the target model, is proposed.

J. Res, Petr Kaska, Martin Perešíni et al. · 0 citations
Open access 2026

Fed-CBE: Client-Side Backdoor Elimination in Federated Learning via Persistent Parameter Disruption

Fed-CBE is proposed, a novel client-side defense algorithm that eliminates backdoors through three synergistic mechanisms: periodic alternating layer resetting disrupts deep parameters to dismantle cross-round backdoor accumulation, and indiscriminate forgetting employs entropy maximization on non-ground-truth classes...

Chun-Hai Li, Yun-Hui Shen, Ming Xie et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Backdoor Containment via Expert Quarantine and Shutdown in LLMs

Backdoored large language models (LLMs) can behave normally on benign inputs while producing attacker-specified outputs under hidden triggers. Existing defenses span four stages--prior-training, in-training, post-training, and inference-time--and share one of two underlying strategies: either suppress backdoor learning...

Jian-Wei Li, Min-Seon Kim, Jung-Eun Kim · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.