SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces
This work presents SkillSafetyBench, a runnable benchmark for evaluating skill-facing safety failures, and suggests that agent safety depends not only on model-level alignment, but also on how agents interpret skills, trust workflow context, and act through executable environments.