Skip to content

ToolFence: Fine-Grained Authorization for Secure Tool-Using LLM Agents

Sep 2026 · 0 citations · 44 references
Computer Science

TL;DR

This work introduces ToolFence, which compiles a typed authorization blueprint before execution, enforces it through a deterministic monitor, and when the blueprint is incomplete asks a judge to grant new capabilities rather than adjudicate each concrete call, improving runtime efficiency.

Abstract

Tool-using LLM agents remain vulnerable to indirect prompt injection because trusted instructions and untrusted observations share one context, allowing malicious content to steer consequential input-filtering defenses. Multi-path consensus defenses still leave a high attack success rate because they examine content or aggregated outputs rather than authorizing effects, especially for the within-tool attack, which preserves the intended tool but manipulates its arguments. Data-Flow Control such as CaMeL provides stronger guarantees, but incurs substantial time latency that limits practical deployment. We introduce ToolFence, which compiles a typed authorization blueprint before execution, enforces it through a deterministic monitor, and when the blueprint is incomplete asks a judge to grant new capabilities rather than adjudicate each concrete call. ToolFence provides two key advantages. First, its fine-grained provenance-aware authorization enables the system to distinguish user-authorized values from untrusted observations, effectively addressing the within-tool attack. Second, its deterministic fast path and capability-level runtime grants substantially reduce the frequency of expensive judge calls, improving runtime efficiency. On AgentDojo with Qwen3-max, ToolFence reduces overall ASR to near zero with only a 3.80 percentage-point clean-utility drop and practical runtime overhead.

View source

Similar papers

Preprint Sep 2026

AGATE: Provenance-Based Runtime Defense Against Compositional Attacks on LLM Agents

LLM agents can produce harmful effects through sequences of ordinary operations. Judging such actions requires establishing both the authority that permits them and the origin of the data they carry. We present AGATE, an authorization and data-provenance gate at instrumented agent-harness boundaries. Operator declarati...

Xiao-Ru Zhang, Zhuo-Ran Cheng, Kai-Lin Liu et al. · 0 citations
#artificial intelligence Review Sep 2026

Zero-Trust Authorization and Discovery for Enterprise MCP

LLM agents translate natural-language context, which may include attacker-controlled text, into privileged tool calls, so authorization must remain effective even when an agent is prompt-injected or adversarially steered. The Model Context Protocol (MCP) has become a widely adopted interface for this boundary, yet its...

Huang-Jian Li, Yu-Wei Wang, Srinivasan Manoharan · 1 citation
Preprint Aug 2026

WebMCP-Phalanx: Enforcing and Characterizing Trust Boundaries for Browser-Integrated LLM Agents

This work proposes WebMCP-Phalanx, a dual-layer agent runtime architecture that provides a browser-native trust anchor that binds each tool to its registering principal through cryptographically protected capability credentials and propagates provenance labels throughout the tool lifecycle.

Lin-Fa Lee, Yi-Yu Chang, Kuo-Hui Yeh · 0 citations
Preprint Oct 2026

Compromise Is Not Consequence: Evaluating Task-Scoped Authorization in LLM Agents with Paired Replay

A tool-using model can follow a malicious instruction even when its credentials are valid. We study whether task-scoped authorization contains the resulting tool execution. Our paired-replay testbed samples a model request once and submits the same action, resource, and arguments to broad bearer, scoped JWT, sender-con...

Tural Hagverdiyev · 0 citations
Preprint Sep 2026

MetaPermit: Scalable and Auditable Access Control for AI Agents via LLM-Inferred Meta-Attributes

The rise of autonomous AI agents equipped with tools has introduced significant security risks, ranging from unintended tool misuse to adversarial manipulation through Indirect Prompt Injection (IPI) attacks. In practice, deployed agent systems such as OpenAI Codex and Claude Code protect tool invocations through a com...

Han-Zhang Ma, Alicia Y. Hariri, Tian-Xiang Shen et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.