Skip to content

OpenSkillRisk: Benchmarking Agent Safety When Using Real-World Risky Third-Party Skills

Jul 2026 · arXiv.org · Vol abs/2607.20121 · 3 citations · ⚡ 1 influential · 44 references
Computer Science

TL;DR

The behavioral analysis reveals three recurring failure patterns: agents may fail to recognize the risk, recognize it but fail to intervene before acting, or follow skill instructions beyond the user's intended scope, which highlights the need to improve both risk reasoning and execution control in agent frameworks.

Abstract

LLM-based agents leverage third-party skills to extend their capabilities in open-world scenarios. However, third-party skills can introduce extra security vulnerabilities, as seemingly harmless skills can contain latent safety risks that only emerge during actual execution. In this work, we conduct a systematic investigation into how well current agent systems recognize and avoid such risks. To support quantitative and qualitative evaluation, we construct OpenSkillRisk, a dedicated safety benchmark containing 263 risky skills collected from public skill marketplaces. We classify these skills into seven categories based on their threat types and pair each skill with a standardized user task and a corresponding sandbox for controlled evaluation. Distinct from prior benchmarks, OpenSkillRisk not only covers more realistic and diverse unsafe scenarios, but also provides a fine-grained analysis to diagnose the behavioral patterns of agents in such scenarios. We conduct comprehensive experiments covering three mainstream CLI agent frameworks and thirteen state-of-the-art LLMs. Experimental results show that no tested system handles risky skills reliably: even the safest configurations still execute unsafe actions in about 17% of cases. Context-dependent and system-level risks are especially difficult for current agent systems to avoid. Our behavioral analysis reveals three recurring failure patterns: agents may fail to recognize the risk, recognize it but fail to intervene before acting, or follow skill instructions beyond the user's intended scope. These findings highlight the need to improve both risk reasoning in LLMs and execution control in agent frameworks.

View source

Similar papers

Preprint Aug 2026

REDAgentBench: Executable Red Teaming and Faithful Measurement of LLM Agent Systems

RedAgentBench is introduced, an executable framework for autonomous red-teaming and faithful measurement that shows that executable evaluation can improve safety measurement and identify actionable intervention points.

Zixing Chen, Xingyuan Liu, Jie Zhu et al. · 3 citations
#artificial intelligence Preprint Sep 2026

Black-Box Red Teaming of Agentic AI: A Taxonomy-Driven Framework for Automated Risk Discovery

Agentic systems are rapidly moving to production, where they read untrusted inputs, call tools with real permissions, and act autonomously, expanding the security surface beyond chat-only models. Yet standard evaluations remain single-turn and fail to capture multi-step agent vulnerabilities. We present a systematic bl...

Divyanshu Kumar, Nitin Aravind Birur, Tanay Baswa et al. · 2 citations
Preprint Aug 2026

How Do LLM Agents Actually Get the Flag? Trace-Level Provenance for Agentic Offensive Security Evaluation

CTF-ABACUS is introduced, a trace-based agent auditing framework that reconstructs each run as an evidence-grounded solve profile that provides a basis for designing benchmarks that better isolate the offensive capabilities of autonomous language-model agents.

Kimberly Milner, Ming-Hao Shao, Nanda Rani et al. · 2 citations
#artificial intelligence Preprint Sep 2026

SINGED: Correct Outputs Do Not Certify Safe Execution in LLM Agents

Tool-using language-model agents select and execute third-party artifacts. Different implementations can return the requested output while producing hidden execution effects that task-, attack-, or choice-based evaluations may miss. We study functional counterfeits: implementations that match benign alternatives on the...

XiaoYu Xu, Zi Liang, Min-Xin Du et al. · 0 citations
#cybersecurity Preprint Aug 2026

The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark

SRE-Bench is introduced, the first realistic, contamination-free RE benchmark, and results indicate that strong source-code security capabilities do not yet transfer to binary analysis, highlighting RE as an important frontier for agentic cybersecurity and SRE-Bench as a rigorous testbed to measure progress.

J. Spence, Nicholas Assaderaghi, Feng Xiao et al. · 1 citation
Preprint Aug 2026

ActBench: Self-Evolving Benchmark of Behavioral Safety in Cowork Agents

Cowork agents may complete benign tasks while disclosing protected data, manipulating unauthorized state, invocate unauthorized API. We define behavioral safety and introduce ActBench, a self-evolving benchmark that evaluates such behavior risk from execution trajectories rather than final responses. Each case pairs a...

H. Yao, Yimin Liu, Meihui Chen et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.