Results show that SAE-based analysis can explain defense fragmentation and guide interpretable backdoor mitigation and show system?atic encoding differences: dirty-label backdoors are dominated by isolated interaction features, whereas clean-label backdoors rely more on heterogeneous mixtures of mixed and weight-modified features.
Abstract
Backdoor attacks pose a serious threat to large language models (LLMs), but existing defenses remain fragmented, failing to pro?vide unified defense against both dirty-label and clean-label attacks. To investigate why such fragmentation arises, we present the first systematic feature-level mechanistic analysis of LLM backdoors using sparse autoencoders (SAEs). Starting from a 2 x 2 comparison of clean and poisoned models on clean and triggered inputs, we trace backdoor-induced logit shifts to high-contributing SAE features and categorize them into four roles: interac?tion, suppressed, mixed, and weight-modified features. This taxonomy reveals system?atic encoding differences: dirty-label back?doors are dominated by isolated interaction features, whereas clean-label backdoors rely more on heterogeneous mixtures of mixed and weight-modified features. These differ?ences explain why existing defenses remain fragmented across attack paradigms. We val?idate this hypothesis through inference-time feature clamping, which reduces ASR to at most 10.8% in most dirty-label settings and at most 15.4% in the majority of clean-label settings, while preserving benign-task perfor?mance. These results show that SAE-based analysis can explain defense fragmentation and guide interpretable backdoor mitigation.
While LLM attacks span prompt optimization, multi-turn context manipulation, retrieval poisoning, and model backdoors, white-box defenses are typically evaluated on isolated attack families. Consequently, whether heterogeneous attacks leave internal representation shifts that generalize to unseen threat mechanisms rema...
Zhen Guo, Shang-Hao Shi, Shamim Yazdani et al.· 0 citations
Bait-and-Recover is proposed, a weight-level defense that places a bait adapter where attackers read activations and a paired recovery adapter at the subsequent layer that decouples the observation path from the behavior path.
Tian Gao, Zhi-Hui Xie, Yu-Hao Wu et al.· 0 citations
AlcaTRAz (Anchored Tree-Rule defense Against jailbreaks), a prompt-level defense based on rule trees that operates exclusively on the input text and requires no modification or retraining of the target model, is proposed.
J. Res, Petr Kaska, Martin Perešíni et al.· 0 citations
A reverse-training framework is introduced that weakens the trigger-target association, producing low-ASR backdoor models while preserving clean-input performance and exposing a fundamental attacker-defender asymmetry in existing defense paradigms.
Fed-CBE is proposed, a novel client-side defense algorithm that eliminates backdoors through three synergistic mechanisms: periodic alternating layer resetting disrupts deep parameters to dismantle cross-round backdoor accumulation, and indiscriminate forgetting employs entropy maximization on non-ground-truth classes...
Chun-Hai Li, Yun-Hui Shen, Ming Xie et al.· IEEE Transactions on Informa...· 0 citations
Backdoored large language models (LLMs) can behave normally on benign inputs while producing attacker-specified outputs under hidden triggers. Existing defenses span four stages--prior-training, in-training, post-training, and inference-time--and share one of two underlying strategies: either suppress backdoor learning...
Jian-Wei Li, Min-Seon Kim, Jung-Eun Kim· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.