Skip to content

Recover, Decode, Reguard: Guard-Agnostic Defense Amplification againstEncoded VLM Jailbreaks

2026 · arXiv.org · Vol abs/2607.26574 · 0 citations · 50 references
Computer Science

TL;DR

An empirical safety– utility ceiling for the non-iterative recovery-based defenses the authors evaluate, recurring across every guard and both target VLMs, is exposed.

View source

Similar papers

#machine learning Preprint Sep 2026

Decoy Direction Optimization: A Post-Hoc Defense Against LLM Abliteration

Decoy Direction Optimization (DDO), a fast, post-hoc weight-editing defense that requires no base-model finetuning, is introduced, based on a simple mechanistic insight: ablation attacks rely on contrastive estimators to find the refusal direction.

Aashiq Muhamed, Mona T. Diab, Virginia Smith · 2 citations · ⚡1
#artificial intelligence Preprint Sep 2026

Decoding Guardrails: XAI-Guided Perturbation Analysis of Prompt Injection Detection

Large language models (LLMs) are increasingly deployed in production systems, raising concerns about their exposure to adversarial manipulation through prompt injection and jailbreak attacks. Classifier-based guardrails, such as Prompt Guard 2, are widely used as a first line of defense against such attacks, but their...

Fernando Outeda, Gustavo Betarte, J. Campo et al. · 0 citations
Preprint Aug 2026

Once Poisoned, Arbitrarily Controlled: A Programmable Backdoor in VLMs

This work shows that a single poisoning phase can implant a programmable backdoor into a VLM, allowing an attacker to choose previously unseen target-caption semantics at inference time and synthesize corresponding stealthy triggers on demand.

Tao Lin, Gao-Jie Jin, Zongxi Liu et al. · 0 citations

Jailbreaking Jailbreaks: A Proactive Defense for LLMs

P RO A CT represents an orthogonal defense strategy that serves as an additional guardrail to enhance LLM safety against the most effective attacks.

Wei-Liang Zhao, Daniel Ben-Levi, Jinjun Peng et al. · 1 citation
Preprint Sep 2026

AlcaTRAz - Anchored Tree-Rule Defense Against Jailbreaks

AlcaTRAz (Anchored Tree-Rule defense Against jailbreaks), a prompt-level defense based on rule trees that operates exclusively on the input text and requires no modification or retraining of the target model, is proposed.

J. Res, Petr Kaska, Martin Perešíni et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.