Autonomous agents increasingly perform long-horizon tasks involving tool use, persistent state, and consequential actions, raising a fundamental question: \emph{under what conditions does an agent cross the boundary of authorized execution while pursuing a legitimate task?} Existing studies often attribute such failure...
Zong-Hao Ying, Xiang-Fan Wu, Bo Yang et al.· 1 citation· ⚡1
Indirect prompt injection embeds malicious instructions within external content retrieved by LLM-based agents, altering target behavior without user authorization. We introduce pikit, a research toolkit designed to systematically evaluate these threats across three core dimensions: attacks (13 methods), channels (16 ca...
Zong-Hao Ying, Xiang-Fan Wu, Bo Yang et al.· 0 citations
How does a multi-agent system evolve from a local deviation into collective loss of control? We propose an epidemic explanation organized around accidental mutation, contagion, and recovery. A spontaneous deviation creates a seed; communication enables other agents to adopt and retransmit its unsafe strategy; collectiv...
Xiang-Fan Wu, Zong-Hao Ying, Hui-Yu Wu et al.· 0 citations
Self-evolving agents increasingly convert interaction histories into reusable skills that persist beyond individual tasks. While prior work studies memory and retrieval poisoning, such attacks only affect agents when poisoned records are retrieved as context. We uncover a new and more fundamental risk: poisoned experie...
Zong-Hao Ying, Xiang-Fan Wu, Hui-Yu Wu et al.· 3 citations
We assess indirect prompt injection in DeepSeek Harness (DSH), using AI-Infra-Guard (A.I.G) to construct tests, deliver controlled taint, execute DSH, collect traces, and judge outcomes. The study covers 14,560 controlled executions over 16 indirect-content channels, text and file carrier modes, 35 payload objectives,...
Zong-Hao Ying, Xiang-Fan Wu, Hui-Yu Wu et al.· 1 citation
This work presents SkillSentry, a dynamic safety-testing framework based on adaptive honey worlds, which infers the intended capability boundary of a skill, constructs an LLM-simulated environment with controlled decoy resources, and adaptively generates tasks to explore its behavioral states.
Nizhang Li, Zong-Hao Ying, Xiang-Fan Wu et al.· 0 citations
AFL and EFL have little detectable route-level association with GPQA-Diamond accuracy and pronounced EFL coincides with a decline in Terminal-Bench pass rate as task exposure increases, a pattern may arise because correctness in long-horizon tasks is more sensitive to extreme fidelity loss.
Xiang-Fan Wu, Zong-Hao Ying, Hui-Yu Wu et al.· 0 citations
SafeFlow is proposed, a defense framework for multi-agent systems that formalizes malicious cross-agent propagation as a semantic information-flow problem and reduces attack success rates compared to undefended baselines and external defenses while retaining high benign task completion and a high paired safe--harm succ...
Haowen Dai, Zonghao Ying, Wenfeng Li et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.