Skip to content

Twin Agent: Context Residual Compression for Privilege Separated Agents

Jul 2026 · arXiv.org · Vol abs/2607.19595 · 0 citations · 26 references
Computer Science

TL;DR

This work proposes Twin Agent, a general privilege separation design pattern inspired by residual coding in the agent context that preserves high task utility while preventing prompt injection attacks, outperforming both undefended agents and privilege separation baselines.

Abstract

Large language model (LLM) agents are vulnerable to security risks, such as prompt injection attacks from untrusted context that manipulate downstream reasoning and tool use. Existing secure-by-design approaches mitigate this risk by separating untrusted observations from privileged execution and careful control of information flow, but often degrade utility and require extensive task-specific engineering. We thus propose Twin Agent, a general privilege separation design pattern inspired by residual coding in the agent context. Twin Agent consists of two nearly symmetric agents: an Explore Agent that inspects untrusted information and a Safe Agent that executes privileged actions. The Explore Agent is conditioned on the Safe Agent's current context and communicates only compact hints to the Safe Agent about the next action to take. This design reduces the information needed to preserve task utility and thus achieves a better security--utility tradeoff, which we empirically verify by measuring how utility and attack success change as the length of hints varies. We evaluate Twin Agent on long-horizon software engineering tasks with SWE-bench Lite and on heterogeneous multi-tool interaction tasks with AgentDojo and DecodingTrust-Agent. Across both benchmarks, Twin Agent preserves high task utility while preventing prompt injection attacks, outperforming both undefended agents and privilege separation baselines.

View source

Similar papers

Preprint Sep 2026

MetaPermit: Scalable and Auditable Access Control for AI Agents via LLM-Inferred Meta-Attributes

The rise of autonomous AI agents equipped with tools has introduced significant security risks, ranging from unintended tool misuse to adversarial manipulation through Indirect Prompt Injection (IPI) attacks. In practice, deployed agent systems such as OpenAI Codex and Claude Code protect tool invocations through a com...

Han-Zhang Ma, Alicia Y. Hariri, Tian-Xiang Shen et al. · 0 citations
Preprint Aug 2026

REDAgentBench: Executable Red Teaming and Faithful Measurement of LLM Agent Systems

RedAgentBench is introduced, an executable framework for autonomous red-teaming and faithful measurement that shows that executable evaluation can improve safety measurement and identify actionable intervention points.

Zixing Chen, Xingyuan Liu, Jie Zhu et al. · 3 citations
Preprint Aug 2026

DreamGuard: Efficient Runtime Guardrail for LLM Agents via Risk-Aware World Model

As large language model (LLM) agents increasingly invoke external tools and interact with real-world systems, unsafe actions may cause irreversible consequences on external states, user data, and downstream services. Recent runtime guardrails mitigate such risks by checking proposed actions before execution, but many r...

Wenhao Lin, Cheng-Yu Yu, Xingwei Lin et al. · 2 citations
Review Aug 2026

ClawSentry: A Progressive Multi-Tier Security Monitor for Safeguarding Autonomous LLM Agents

This work argues that agentic risk is progressive: it can enter at four loci of the agent control loop--skill admission, invocation-time intent, execution-time effect, and post-action consequence--while a denied dangerous objective can reappear across surface forms, tools, or turns.

Kai Wang, Zeming Wei, Biaojie Zeng et al. · 0 citations
Review Sep 2026

Separation of Duties for Privileged LLM Agents: A Governed Execution Architecture with Measured Security-Utility Trade-offs

Large language model agents are increasingly granted real privileges (executing commands, modifying files, calling APIs), so an agent that errs has already acted. Existing defences concentrate on the agent's inputs, while the path from a candidate action to privileged side effects remains less directly studied. We argu...

Qi-Shuai Jing · 0 citations
Review Aug 2026

Compositional Threat Analysis of Latent Compromise in LLM Agent Systems: The Order 66 Scenario

This paper introduces a compositional model explaining why no component is catastrophic alone, yet their conjunction can produce correlated destructive action, and separates three population-reach routes from a common core of dormancy, activation, authority, reachable targets, and failed recovery.

Satoshi Matsuoka · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.