LLM agents can produce harmful effects through sequences of ordinary operations. Judging such actions requires establishing both the authority that permits them and the origin of the data they carry. We present AGATE, an authorization and data-provenance gate at instrumented agent-harness boundaries. Operator declarations and host approval events ground authorization; delegated actions are constrained by grants that bind to exact parameters, expire, and permit a limited number of uses. Source registration connects observed inputs to subsequent transfers, while an effect ledger tracks repeated requests. Deterministic checks make decisions without an LLM in the decision path and retain their grounds with execution evidence for forensic replay. Adapters integrate three production harnesses -- DeepSeek Harness, OpenCode, and OpenClaw -- without modifying host code, translating each host's native observation and veto points into a single shared gate interface; the judgment core is identical in all three, and only enforcement depth differs. Our evaluation combines 153 exercised attack-chain records with deployment, utility, and reconstruction experiments. The deployment observations expose how tool declarations and data checks govern business actions, including a bypass through parameter rewriting. Six of eleven benign file-processing scenarios contain denial events, revealing the utility cost of content-based provenance policies. Across 252 runs on 63 sanitized scenarios, replay agrees with live graph projections for all 63 scenarios on each of two platforms. These results establish the feasibility of provenance-based runtime judgment and identify content transformation, legitimate reuse, and observation coverage as concrete limits.
This work introduces ToolFence, which compiles a typed authorization blueprint before execution, enforces it through a deterministic monitor, and when the blueprint is incomplete asks a judge to grant new capabilities rather than adjudicate each concrete call, improving runtime efficiency.
Yan-Jie Li, Xiang-Yu He, Xue-Long Dai et al.· 0 citations
Tool-using large language model (LLM) agents turn generated text into real side effects, so poisoned tool metadata, retrieved pages, memory, and reusable skills can steer the next call. Vetting an artifact before admission does not settle this. A safe variant and a leaking variant can produce the same admission evidenc...
Feng-Peng Li, Qi-Zhou Wang, Yu-Ke Hu et al.· 0 citations
SkillGuard is presented, a harness-level enforcement layer that treats this event as contamination and restricts future capabilities to disconnect the resulting state from deployer-defined forbidden states and preserves substantially more capabilities than binary restriction at the same attack success rate.
Wu-Jie Xiong, Rabimba Karanjai, Yang Lu et al.· 2 citations
KITA is presented, a review-to-authorization architecture that keeps the user's personal secret signing key and every threshold signing-key share outside all LLM processes and establishes execution-bound authorization integrity.
This work proposes SARA, which treats action induction and execution authorization as distinct runtime roles and separates action provenance from execution authority, and applies No-History-Promotion to prevent historical recurrence from laundering action origins into execution authority.
Xiao-Kun Guo, Zhen Xu, Dongdong Huo et al.· 0 citations
The rise of autonomous AI agents equipped with tools has introduced significant security risks, ranging from unintended tool misuse to adversarial manipulation through Indirect Prompt Injection (IPI) attacks. In practice, deployed agent systems such as OpenAI Codex and Claude Code protect tool invocations through a com...
Han-Zhang Ma, Alicia Y. Hariri, Tian-Xiang Shen et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.