Self-evolving skill (SES) systems distill agent trajectories into persistent skills, allowing untrusted experience to become trusted instruction. We introduce PoisonedEvolution, a trajectory-poisoning attack on this promotion process. Our skill-visible black-box attacker can inspect a target skill and contribute bounde...
Jia-Luo Chen, Lingqi Jiang, Xin-Hao Deng et al.· 0 citations
Computer-use agents increasingly interact with browsers, terminals, file systems, and external services, introducing safety risks that emerge through runtime behavior rather than generated content alone. Existing guard models target static prompts and responses and are poorly suited to agent execution; existing executa...
Yun-Hao Feng, Rui-Xiao Lin, Ming Wen et al.· 0 citations
SKILLTRACE is presented, a multi-trace provenance auditing framework for LLM-agent skill reuse that represents the Operational Trace as a Skill Operational Graph (SOG) that captures activation, procedure, and resource-flow structure.
Jia-Luo Chen, Minghe Wang, Lingqi Jiang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.