OpenSkillRisk: Benchmarking Agent Safety When Using Real-World Risky Third-Party Skills
The behavioral analysis reveals three recurring failure patterns: agents may fail to recognize the risk, recognize it but fail to intervene before acting, or follow skill instructions beyond the user's intended scope, which highlights the need to improve both risk reasoning and execution control in agent frameworks.