A reverse-training framework is introduced that weakens the trigger-target association, producing low-ASR backdoor models while preserving clean-input performance and exposing a fundamental attacker-defender asymmetry in existing defense paradigms.
Abstract
Backdoor attacks are among the most effective and stealthy attacks in deep learning. Existing attacks and defenses are largely designed and evaluated under the assumption that successful backdoors exhibit high Attack Success Rates (ASRs). In this paper, we show that this assumption creates a fundamental weakness in existing defense paradigms. ASR is not an intrinsic property of a backdoor; rather, it is an attacker-controlled variable that can be deliberately reduced without eliminating the underlying backdoor behavior. We introduce a reverse-training framework that weakens the trigger-target association, producing low-ASR backdoor models while preserving clean-input performance. Through extensive evaluation across multiple datasets, diverse attack families, and multiple architectures, we show that state-of-the-art defenses fail consistently under low-ASR conditions, exposing a fundamental attacker-defender asymmetry.
Deep neural networks (DNNs) are vulnerable to backdoor/Trojan attacks, posing a severe threat to security-critical applications. In this paper, we investigate a scenario in which two different backdoors are successively embedded into one model and report that the later-embedded backdoor has a “suppression” effect on th...
Hong Zhu, Kai Chen, Sheng-Zhi Zhang· IEEE Transactions on Informa...· 0 citations
Problem-space evasion attacks have exposed critical weaknesses in machine learning-based malware detectors; yet, their evaluation remains fragmented across models, datasets, and attack methodologies, often neglecting domain-specific requirements such as executability and functionality preservation. We address this gap...
Mashal Zainab, Salijona Dyrmishi, Hamid Bostani et al.· 0 citations
RAMP is proposed, an attack enhancement method that uses a genetic algorithm to optimize reversed adversarial perturbations under black-box access and then injects them through functionality-preserving binary manipulations, which substantially improves attack effectiveness over trigger-only baselines.
Jin-Wen Xin, Dong-Ni Zhang, Chen-Yang Wang et al.· 0 citations
Results show that SAE-based analysis can explain defense fragmentation and guide interpretable backdoor mitigation and show system?atic encoding differences: dirty-label backdoors are dominated by isolated interaction features, whereas clean-label backdoors rely more on heterogeneous mixtures of mixed and weight-modifi...
Yikun Zeng, Chenxu Niu, Wei Zhang et al.· 0 citations
Dueling bandit algorithms excel in learning from pairwise comparisons, offering robust performance guarantees in benign environments. However, recent evidence suggests that even state-of-the-art methods can be highly susceptible to adversarial manipulation. In this work, we introduce and analyze a post-action attack mo...
Mo Lyu, Chen-Ye Yang, Guan-Lin Liu et al.· IEEE Transactions on Signal...· 0 citations
BackDFL is presented, a unified benchmark for systematically evaluating DFL under realistic and adaptive backdoor attacks, and demonstrates that both state-of-the-art Byzantine-robust DFL methods and adapted FL backdoor defenses fail under modest malicious participation rates, especially in heterogeneous settings.
M. Bouchiha, Gregory Blanc, Yu-Fei Han· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.