Preprint
Aug 2026
A Self-Evolving Multi-Agent Framework Defense against LLM Jailbreak Attacks
This work proposes a self-evolving test-time defense built around a persistent, cross-interaction rule memory that substantially reduces attack success rates while preserving benign utility, remains robust under an adaptive composite-wrapper attack, and does not increase over-refusal as the memory grows.
Tongshen Hu, Bryan Hooi
· 0 citations