Skip to content

SkillDRE: Dual-Stage Red-Team Evolution of Agent Skills via Pre-Execution and Runtime Feedback

Sep 2026 · 0 citations · 23 references
Computer Science

TL;DR

Results show that two-stage defense feedback can serve as a useful learning signal for adaptive red teaming and that evaluating either defense stage in isolation can miss the resulting attack capability.

Abstract

Agent skills package instructions, executable code, and task-specific resources into reusable artifacts that agents can improve using execution feedback. The same mechanism also enables attackers to evolve malicious skills, making them more effective and less detectable. However, a candidate skill may pass pre-execution scanning yet fail to realize its target under runtime defenses, while a revision that repairs execution may introduce new scanner findings. We introduce SkillDRE, a fully automated framework for evolving complete malicious skill packages through a dual-stage feedback loop. Given a benign task and its associated skills, SkillDRE autonomously constructs and validates a task-conditioned malicious objective and a verifiable judge rule. It then holds both fixed while evolving the skill implementation, with preservation of legitimate task capability. SkillDRE combines scanner-guided evolution with runtime-guided refinement informed by execution outcomes observed under runtime defense. Each runtime-guided revision returns to the pre-execution stage for rescanning and further optimization before re-execution, forming a cross-stage closed loop. Evaluated on SkillsBench across four victim models, SkillDRE achieves an average attack success rate of 45.28%, exceeding the strongest baseline by 40.3%, while its final submitted skills receive no SkillScan findings and largely preserve benign-task performance. These results show that two-stage defense feedback can serve as a useful learning signal for adaptive red teaming and that evaluating either defense stage in isolation can miss the resulting attack capability. Codes is available at https://github.com/whfeLingYu/SkillDRE

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Defense-as-Skill: Evolving Runtime Guard Skill for Skill-Augmented Agents

This work proposes Defense-as-Skill, a defense paradigm that implements the runtime guard itself as an installable, inspectable, and editable skill, and demonstrates transfer across victim models, held-out risk families, and external benchmarks, as well as retained protection against adaptive attackers.

Xiao-Fan Yang, Ziqi Miao, Dian-Bo Sui et al. · 0 citations
Preprint Aug 2026

TRUSS: Towards Task-Reliable and User-Safe Automated Agent Skill Generation

Results show that execution evidence can expose behavioral failures missed by artifact inspection and can guide Skill generation toward jointly verified functional and safety outcomes.

Zhi-Bo Zhang, Zheng-Mao Ouyang, Ling Shi et al. · 0 citations
Preprint Aug 2026

SkillSentry: Reliable Skill Execution for LLM Agents via Runtime Assurance

SkillSentry is proposed, a skill-oriented runtime assurance framework built upon a new domain-specific language (DSL) for representing runtime guidance for skill execution that improves the task success rate of LLM agents by 24.1% across skills, on average, while exhibiting lower variability across repeated runs.

You Lu, Xinyu Huang, Bi-Huan Chen et al. · 1 citation
Preprint Aug 2026

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution

LLM agents increasingly execute long-horizon tasks through tool use and environment interaction, shifting evaluation from final-response scoring to verification of complete executions. For skill-augmented agents, verification additionally requires the procedural knowledge encoded in task-time skills, because this knowl...

Zhien Han, C. Zeng, Liuhaichen Yang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Subagents vs Agent Skills: Executing Reusable Knowledge for Long-Horizon Agentic Tasks

How can language model agents effectively leverage libraries of reusable knowledge to solve long-horizon tasks? Recent work has increasingly focused on agent skills: reusable capabilities represented as skill packages, i.e., multi-file bundles containing instructions, scripts, and other resources that help agents perfo...

Wasu Top Piriyakulkij, Rachel Lawrence, Alicia Curth et al. · 0 citations
Preprint Aug 2026

From Relevance to Execution Utility: Reward-Aware Dynamic Execution Gating for Skill-Based LLM Agents

Agent skills are increasingly used to equip large language model (LLM) agents with reusable procedural knowledge. Although recent work has substantially improved skill retrieval due to the increasing skill libraries, retrieving a plausible skill bundle does not guarantee that executing it is worthwhile. Since every ski...

Liang He, Jing Wen, Hong-Yu Gu et al. · 2 citations · ⚡1

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.