It is found that identical safety feedback induces sharply model-dependent behavioral routing rather than uniform protection, establishing a mechanistic lens and a predictive baseline for anticipating the safety and utility trade-offs of post-jailbreak feedback in autonomous agents.
Abstract
As large language models increasingly operate as tool-using agents, post-jailbreak safety feedback is often assumed to serve as a reliable safeguard; however, how lingering jailbreak context shapes subsequent agent behavior remains largely unexplored. To systematically examine this dynamic, we introduce a paired continuation framework across 192 parent tasks spanning 42 domains, evaluating 12,148 analyzed continuation pairs (curated from a 12,288-pair initially design) across eight diverse agents. We find that identical safety feedback induces sharply model-dependent behavioral routing rather than uniform protection: redirecting unsafe trajectories toward legitimate completion (\emph{rescue}), sustaining unauthorized execution (\emph{persistent unsafe}), or triggering over-refusal on benign tasks (\emph{collateral loss}). Through layer-wise activation patching, we discover a shared \emph{late-commit pattern} where causal intervention effects surge sharply near the final layers (relative depths of 0.958--0.984) despite an over 30-fold variation in peak magnitude across architectures. Crucially, critical-layer representations correlate with macroscopic routing outcomes, and intervening at these layers causally alters concrete next-step tool actions. Building on this causal foundation, we test whether localized intervention-derived features can serve as predictive proxies for full-trajectory routing outcomes on unseen parent tasks under leave-one-parent-task-out evaluation, finding that they provide viable predictive signals in responsive agents with peak ROC AUCs reaching 0.675 for \emph{rescue}, 0.777 for \emph{collateral loss}, and 0.702 for \emph{persistent unsafe}. These findings establish a mechanistic lens and a predictive baseline for anticipating the safety and utility trade-offs of post-jailbreak feedback in autonomous agents.
Convergent Detour Hijacking is introduced, a text-only, runtime-independent attack that couples selection manipulation, malicious skill instructions, and tool-chain resource amplification largely separately and evaluates it across multiple LLM backends and 491 held-out tasks under single-task and multi-turn conditions.
Junliang Liu, Ruo-Yu Li, Wenxin Tang et al.· 1 citation
The transition from Large Language Models (LLMs) to agents shifts safety stakes from toxic text to irreversible environmental harm. While current defenses remain largely retrospective, proactive runtime intervention is bottlenecked by the lack of large-scale, causally-consistent data. We propose PROACT-Agent, a framewo...
Ding Jia, Wei Liu, Xiang-Long Du et al.· 0 citations
MazeRunner is presented, an autonomous penetration testing system built on a three-agent task-and-clue orchestration framework that separates global orchestration, context-intensive execution, and failure-oriented review while persistently maintaining task states and environmental evidence.
Zhenyuan Li, Yi Jiang, Junjie Cheng et al.· 0 citations
This study introduces SymTrace, a controlled evaluation framework that records the MAS execution trajectory and establishes intervention anchors and explores the effectiveness of MAS repair methods, revealing that existing unguided rerun methods are highly unreliable.
Zhong-Wen Luan, Xiaoyan Zhang, Ming Hu et al.· 2 citations
Cowork agents may complete benign tasks while disclosing protected data, manipulating unauthorized state, invocate unauthorized API. We define behavioral safety and introduce ActBench, a self-evolving benchmark that evaluates such behavior risk from execution trajectories rather than final responses. Each case pairs a...
H. Yao, Yimin Liu, Meihui Chen et al.· 0 citations
Large language model (LLM)-powered agents can be accurate on average yet unreliable in production, a discrepancy that has been observed but remains largely unaddressed. When given the same task five times, a ReAct agent on the AppWorld benchmark using GPT-4.1 succeeds in all five runs only 53% of the time, even though...
Evelyn Duesterwald, Benjamin Elder, Lilian Ngweta et al.· 0 citations
With $2.1 million funding from Google.org, the open-source Public Transit Intelligence Hub will unify public transit monitoring, operations, and passenger communication.
Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.