Skip to content
Preprint

Benign Alone, Harmful Together: Exploiting Experience Composition in Self-Evolving LLM Agents

Aug 2026 · 2 citations
Computer Science

TL;DR

EvoBreak is an experience-conditioned sequential attack that operates through individually benign attack-stage tasks and induced experiences, revealing benign experience composition as a persistent attack surface in self-evolving agents.

Abstract

Self-evolving large language model agents improve their capabilities by distilling interaction trajectories into persistent experiences. Yet this mechanism introduces a new safety risk: experiences that are benign in isolation may jointly weaken an agent's safety boundary when accumulated and reused across sessions. Existing memory attacks typically require direct memory access or induce explicitly malicious records, limiting their stealthiness and applicability. We propose EvoBreak, an experience-conditioned sequential attack that operates through individually benign attack-stage tasks and induced experiences. EvoBreak repeatedly observes the experiences distilled by the victim, identifies uncovered target-relevant requirements, and adaptively acquires complementary experiences before reformulating the final query to activate them jointly. To support training, we introduce BreakGym, a structure-first synthesis pipeline that generates decomposable safety-sensitive targets with diverse dependency structures. EvoBreak is optimized using rejection-sampling supervised fine-tuning and Hint-guided GRPO. Experiments across self-evolving frameworks, victim backbones, pre-evolution domains, and safety benchmarks demonstrate that EvoBreak consistently outperforms existing attacks while maintaining high benignness. These results reveal benign experience composition as a persistent attack surface in self-evolving agents.

View source

Similar papers

Preprint Aug 2026

SynChain: Inducing Computer-Use Agent Systems to Construct Their Own Attack Chains

This work introduces SynChain, a self-synthesized attack paradigm utilizing persistence-aware directed supervised fine-tuning to induce agents to create poisoned yet benign-looking artifacts, proving that securing CUAs requires provenance-aware reasoning over cross-task execution trajectories.

Fu-Yao Zhang, Jia-Ming Zhang, Che Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SEABench: Benchmarking Endogenous Misalignment In Self-Evolving Agents

Self-evolving LLM agents have gained prominence for their ability to improve after deployment by modifying their harness, including their controller instructions, memory management protocols, and reusable tools and skills, in response to user and environment feedback. However, locally useful updates may persist into la...

Saswat Das, Parvati Viswanathan, Daniel Donnelly et al. · 0 citations
Preprint Sep 2026

Self-Evolving Defense: Continual Security Policy Learning for LLM Agents

Large language models (LLMs) increasingly power agents that access sensitive information, use external tools, and modify software repositories. Although these capabilities offer substantial benefits, they also create security risks such as jailbreaks, prompt injection, and vulnerable code generation. Existing defenses...

Minh Nhat Le, Nisarga Gondi, Yi-Bo Peng et al. · 0 citations
Preprint Aug 2026

Query-Only Backdoor Attacks on Self-Evolving Skills via Trajectory Poisoning

Agentic skills improve large language model (LLM) agents by encoding reusable procedures for complex tasks. However, manually authored skills often adapt poorly to long-horizon tasks and changing environments. To address the limitation, self-evolving skill systems have been developed to automatically construct and upda...

Yuyang Luo, Haoran Wang, Kai Shu · 0 citations
Preprint Aug 2026

A Self-Evolving Multi-Agent Framework Defense against LLM Jailbreak Attacks

This work proposes a self-evolving test-time defense built around a persistent, cross-interaction rule memory that substantially reduces attack success rates while preserving benign utility, remains robust under an adaptive composite-wrapper attack, and does not increase over-refusal as the memory grows.

Tongshen Hu, Bryan Hooi · 0 citations
#artificial intelligence Preprint Sep 2026

CoDeL: Co-Evolutionary Defense against Indirect Prompt Injection in LLM-based Agents

Large language model (LLM)-based agents increasingly rely on external tools and content, exposing them to indirect prompt injection (IPI). This threat has motivated a wide range of defenses, among which training-based defenses are often regarded as most reliable. However, existing training-based defenses are typically...

Xiao Yang, Yang-Chen Ou, Yu-Han Gao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.