Skip to content
Preprint

AlcaTRAz - Anchored Tree-Rule Defense Against Jailbreaks

Sep 2026 · 0 citations · 47 references
Computer Science

TL;DR

AlcaTRAz (Anchored Tree-Rule defense Against jailbreaks), a prompt-level defense based on rule trees that operates exclusively on the input text and requires no modification or retraining of the target model, is proposed.

Abstract

Large language models (LLMs) are vulnerable to jailbreak attacks that bypass safety alignment through carefully crafted prompts. Many existing defenses require access to model weights or internals, making them difficult to apply to black-box deployments. We propose AlcaTRAz (Anchored Tree-Rule defense Against jailbreaks), a prompt-level defense based on rule trees that operates exclusively on the input text and requires no modification or retraining of the target model. The method automatically learns a transferable transformation rule that inserts controlled character-level perturbations at selected positions, thereby disrupting structural regularities exploited by jailbreak attacks while largely preserving the model's utility on benign queries. We evaluate the proposed method across 33 open-weight models, 22 jailbreak attack types, and a benchmark of short, single-turn benign questions, comparing against three representative prompt-level baselines (Llama Guard, RA-LLM, Goal Prioritization). Among the compared defenses, AlcaTRAz achieves the best composite security and functionality score in 73.4 % of model-attack combinations and shifts the aggregate score from a modal value of 10 (maximal-severity response to the malicious request) in the undefended setting to a modal value of 2 (near-refusal) after defense, while keeping the mean benign score within 0.27 points of the undefended baseline (8.35 vs. 8.62 on a 0-10 scale). AlcaTRAz substantially reduces but does not eliminate jailbreak success: a high-severity tail remains, and we do not consider adaptive attackers, so we position it as one layer within a defense-in-depth strategy rather than a standalone guarantee.

View source

Similar papers

Jailbreaking Jailbreaks: A Proactive Defense for LLMs

P RO A CT represents an orthogonal defense strategy that serves as an additional guardrail to enhance LLM safety against the most effective attacks.

Wei-Liang Zhao, Daniel Ben-Levi, Jinjun Peng et al. · 1 citation
2026

Semantic Isomorphism Attacks and Defense Evaluation for Jailbreaking Large Language Models

The safety alignment of large language models (LLMs) faces persistent challenges from jailbreak attacks. While existing methods mostly leverage prompt engineering or adversarial optimization, we identify and formalize an underexplored semantic isomorphism vulnerability where harmful and safe scenarios share highly cons...

Fan Yang, Ke Wang, Wenzhou Dou et al. · 0 citations
#natural language process... Preprint Sep 2026

CASCADE Against Jailbreaks: Combination Across Stages with Controlled Attack-Defense Evaluation

Defenses against jailbreak attacks on Large Language Models (LLMs) operate at different pipeline stages, such as input modification or output guard, but it remains unclear which defenses to deploy at each stage and how to combine them. Prior empirical studies, fragmented by inconsistent attack-success-rate definitions...

Jia-Le Luo, Eric Han · 1 citation
Conference Open access Sep 2026

Explaining Jailbreaks: Structured and Interpretable Safety Assessment for Large Language Models

This work proposes an explanation-aware safety framework that augments binary harmfulness detection with structured, human-interpretable explanations capturing severity, strategies, trigger spans, ratio-nales, and derived safety factors, and introduces a human–LLM hybrid annotation and canonicaliza-tion pipeline.

Sunghee Dong, Sungwon Yi, K. Bae et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.