It is revealed that Bit-Flip Attacks (BFAs) can serve as an attack vector for inducing decision-level hijacking, requiring no real-time interaction or control over the training process, and only a minimal number of weight bits need to be flipped after deployment to achieve stealthy, low-cost, and persistent cognitive manipulation.
Abstract
Large Language Models (LLMs) have been widely applied in high-stakes decision-making scenarios such as corporate strategy, and users are increasingly relying on their outputs. However, the deep integration of open-source model sharing ecosystems with LLM-powered critical decision-making applications also introduces critical risks: if an attacker can manipulate the model's cognitive stance, they can indirectly influence the judgments and actions of downstream decision-makers. This paper defines such threats as decision-level hijacking. Existing attacks fail to achieve targeted cognitive manipulation without triggering prohibited content or degrading model functionality. To fill this gap, this paper reveals that Bit-Flip Attacks (BFAs) can serve as an attack vector for inducing decision-level hijacking, requiring no real-time interaction or control over the training process, and only a minimal number of weight bits need to be flipped after deployment to achieve stealthy, low-cost, and persistent cognitive manipulation. Therefore, we propose CogBias, a cognitive bias injection framework for LLMs. CogBias converts subjective preferences into optimization signals via a differentiable sentiment evaluator, uses a multi-objective loss to jointly constrain multiple dimensions, and constructs BitScout to locate critical bits, achieving targeted cognitive intervention under an ultra-sparse flip budget. Experiments on Llama-3.2-3B, Mistral-7B, and Qwen2.5-14B, as well as on the commercial recommendation and controversial factual topic scenarios, demonstrate that flipping only a small number of bits stably induces significant stance shifts on target topics, while the impact on non-target tasks and overall output distribution is limited. This work demonstrates that minute perturbations to low-level weight data suffice to undermine the high-level value alignment of LLMs.
A red-teaming testing method for fine-tuning-stage defenses that audit the harmful content of poisoned fine-tuning data, exposing a vulnerability of the RFT data supply chain to logic injection and point to the need for fine-tuning-stage defenses that audit the harmful content of poisoned fine-tuning data.
Qiuxiang Li, Ke Xu, Yu-Bin Qu et al.· International Conference on...· 0 citations
Mixture-of-Experts (MoE) architectures enable scalable and efficient large language models (LLMs) by selectively activating expert sub-networks through a routing mechanism. However, this adaptive design introduces a new attack surface: specific experts become disproportionately correlated with certain tokens (e.g., end...
Hua-Kang Lin, Tian-Cheng Zheng, Ming-Xuan Sun et al.· 0 citations
Large language models are often fine-tuned, shared, or downloaded from third parties, so a deployed model may carry a hidden backdoor that behaves normally on benign inputs but switches to attacker-controlled behavior when a secret trigger appears. While backdoors can be audited before deployment, runtime monitoring re...
Wen Rui, Ahmed Salem, Andrew Paverd et al.· 0 citations
AEGIS (Adaptive Ensemble Guard for Injection Shielding) extracts instruction-sensitive projectors to identify malicious instructions and leverages a Unified Multi-Layer Consensus mechanism that aggregates topologically distinct signals across the network depth.
Jia-Hao Chen, Ruiping Yin, Xin-Feng Li et al.· 1 citation· ⚡1
A defense taxonomy spanning three axes, namely prompt-level, inference-time, and training-time interventions, is proposed, within which 30 mitigation mechanisms published from 2024 onwards are systematically analyzed, demonstrating that no single defense mechanism provides comprehensive protection, and that robust depl...
Berkay Özçam, Mustafa Kara, Muhammet Ali Aydin et al.· Electronics· 0 citations
The increasing integration of large language models (LLMs) into systems introduces new attack surfaces that extend beyond traditional software vulnerabilities. While LLMs are commonly protected by prompt-level security mechanisms, recent researches show that these controls can be bypassed through carefully crafted inpu...
Tamás Girászi, Natália Papp, Norbert Oláh et al.· Proceedings of the 13th Inte...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.