Skip to content
Preprint

$S^3$: Improving Agent Safety through Multi-Stage Defense

Aug 2026 · 0 citations · 20 references
Computer Science

TL;DR

An automated transformation pipeline that converts existing safety designs into reusable safety skills and establish a community-driven safety skill library is developed and results demonstrate the potential of stage-specific safety skills as a scalable and composable foundation for building resilient and trustworthy agent systems.

Abstract

Large Language Model (LLM) agents rely on multi-stage agentic workflows, with stages such as memory, planning, and tool execution, to accomplish complex tasks. However, risks may emerge at different stages, propagate across steps, and become difficult to detect and mitigate. Existing safety methods protect only isolated stages and are difficult to integrate, leaving agents without comprehensive protection throughout the workflow. To address these limitations, we introduce Stage-Specific Safety Skills, a unified abstraction that represents heterogeneous safety designs as reusable and composable components with explicit stage semantics. We further develop an automated transformation pipeline that converts existing safety designs into reusable safety skills and establish a community-driven safety skill library. Building on this abstraction, we propose $S^3$, a multi-stage defense framework in which a guard agent orchestrates stage-specific safety skills for risk detection and mitigation throughout the agentic workflow. We also construct the Multi-Stage Risk Benchmark (MSRB) to evaluate representative risks across workflow stages. Experimental results show that $S^3$ consistently outperforms representative state-of-the-art baselines in both safety effectiveness and utility preservation. These results demonstrate the potential of stage-specific safety skills as a scalable and composable foundation for building resilient and trustworthy agent systems.

View source

Similar papers

Review Open access Aug 2026

A Systematic Survey of Agentic Skills: Architecture, Lifecycle, and Security

Autonomous large language model (LLM) agents increasingly face reliability, context consumption, and execution stability bottlenecks when deployed on complex, long-horizon tasks. While monolithic prompt engineering and stateless tool-calling paradigms struggle to scale, the field is rapidly converging toward \emph{agen...

Sanket Badhe, D. Shah, Priyanka Tiwari et al. · 0 citations
#natural language process... Preprint Sep 2026

MTAC-IFBench: Benchmarking Instruction-Following in Multi-Turn Agentic Coding

Recently, the rapid development of large language models (LLMs) has reshaped software engineering by enabling autonomous code agents that plan, execute, and utilize external tools iteratively to tackle complex tasks. Beyond achieving functional correctness, these agents must faithfully follow process instructions and c...

Bo-Si Wen, Cun-Xiang Wang, Jia-Yi Gui et al. · 0 citations
Preprint Aug 2026

SHE: Trajectory-driven Safety Harness Evolution for LLM Agents

Safety Harness Evolution (SHE) is proposed, a framework that learns evolving safe boundaries from rollout trajectories and introduces an attribution-guided evolution loop that converts trajectory failures into structured diagnoses, learns artifact-specific boundary refinements, and selects evolved harnesses through saf...

Wanying Qu, Qing-Hua Mao, Yu Li et al. · 5 citations
#artificial intelligence Preprint Aug 2026

ComponentBench: Diagnosing Component-Level Failures in Computer-Use Agents

This work presents ComponentBench, a benchmark and diagnostic pipeline for component-level evaluation of computer-use agents on modern web UIs, and introduces a scalable pipeline for auditing realized structural difficulty after implementation and synthesizing structured failure analyses across tasks and component fami...

Tian-Chen Guan, Xinlei Lin, Royce Cheng-Yue et al. · 0 citations
Open access Sep 2026

Security-Augmented Use Case Flow Refinement for LLM-Based Agentic Systems

Unlike conventional software systems and large language models (LLMs) evaluated in isolation, LLM-based agentic systems introduce a compositional, workflow-level attack surface. Existing security analysis approaches provide limited support for identifying security omissions implicit in functional use case flows and der...

Guang-Yu Wang, Bangqi Li, Ji Wu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.