An empirical safety– utility ceiling for the non-iterative recovery-based defenses the authors evaluate, recurring across every guard and both target VLMs, is exposed.
Decoy Direction Optimization (DDO), a fast, post-hoc weight-editing defense that requires no base-model finetuning, is introduced, based on a simple mechanistic insight: ablation attacks rely on contrastive estimators to find the refusal direction.
Aashiq Muhamed, Mona T. Diab, Virginia Smith· 2 citations· ⚡1
Bait-and-Recover is proposed, a weight-level defense that places a bait adapter where attackers read activations and a paired recovery adapter at the subsequent layer that decouples the observation path from the behavior path.
Tian Gao, Zhi-Hui Xie, Yu-Hao Wu et al.· 0 citations
Large language models (LLMs) are increasingly deployed in production systems, raising concerns about their exposure to adversarial manipulation through prompt injection and jailbreak attacks. Classifier-based guardrails, such as Prompt Guard 2, are widely used as a first line of defense against such attacks, but their...
Fernando Outeda, Gustavo Betarte, J. Campo et al.· 0 citations
This work shows that a single poisoning phase can implant a programmable backdoor into a VLM, allowing an attacker to choose previously unseen target-caption semantics at inference time and synthesize corresponding stealthy triggers on demand.
Tao Lin, Gao-Jie Jin, Zongxi Liu et al.· 0 citations
AlcaTRAz (Anchored Tree-Rule defense Against jailbreaks), a prompt-level defense based on rule trees that operates exclusively on the input text and requires no modification or retraining of the target model, is proposed.
J. Res, Petr Kaska, Martin Perešíni et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.