Skip to content

DGF-Bench: A Benchmark for Simulating and Auditing Deception Against Multi-Agent Governance Boards

Sep 2026 · 0 citations
Computer Science

TL;DR

DGF-Bench is a benchmark in which a board of agents (specialist gates and a General gate that consolidates their decisions) reviews synthetic dossiers while an attacker plants deceptive content in evidence the organization does not vouch for.

Abstract

Tool-using language-model agents can review enterprise projects as governance boards do: they read the evidence, apply written rules and decide whether the project may proceed. Part of that evidence comes from suppliers and project members with a stake in the decision. DGF-Bench is a benchmark in which a board of agents (specialist gates and a General gate that consolidates their decisions) reviews synthetic dossiers while an attacker plants deceptive content in evidence the organization does not vouch for. Dossiers are generated from canonical facts under 61 executable rules, with 42 authoritative records and 32 narrative documents; every gate is certified decidable from those records. Attacks never change an authoritative value, so an attacked dossier keeps the reference decisions of its clean copy. A success is attributable only when the agent receives the injection and takes the exact injected action, which it does not take on the paired clean dossier; the DGF score is the share of applicable fixed attacks a model blocks. Reading documents and records themselves, five of six models were outcome-strict (disposition, findings, actions and authorization all correct) on 82 to 85 of 85 gates. Over 2,622 attacked gate runs, seven direct-order, false-data and false-authority attacks obtained one attributable success against these five, whereas task-aligned attacks imitating the organization's own process passed against four of them: a record note citing a fake review procedure lowered GPT-6 Luna Pro from 34 to 6 outcome-strict gates and DeepSeek V4 Pro from 33 to 7. DGF scores ranged from 96.2 to 26.9, and a policy-aware adaptive attacker writing in records succeeded against five of six models. The approval tool executed no forged approval, yet deceived agents submitted approvals that the rules forbid. The open-source package dgf-bench computes the DGF score with one command.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Persuaded, Not Informed: Incentive-Misaligned Witnesses Defeat In-Context Grounding

This work identifies a failure mode not addressed by a stronger model: when the context contains an assertion by a party with an incentive toward optimism - here the sales representative, a witness recorded in the CRM - the model treats the assertion as evidence and clears deals the company's own records deem unaccepta...

Rahul Balakavi · 0 citations
Preprint Oct 2026

What a Policy Gate Can and Cannot Know: Measured Boundaries of Cross-Platform Command Adjudication

Gateways that adjudicate an agent's actions before they execute are only as good as their understanding of the action. We study a policy gate that never parses shell syntax: it consumes a typed, realised action (verb, operands, resolved zones, program-object identity) and decides ALLOW, ASK or DENY. Working on a Linux...

Qi-Shuai Jing · 0 citations
#artificial intelligence Preprint Oct 2026

Verify Claims, Not Scores: Evidence-Based Verification of Modular Agents

When developers change one component of an agent, such as its controller, a learned model or its verifier, they usually judge the change by an aggregate task score. That score cannot tell whether improvement was attainable, which component lost value, or what the agent's own checks certify. We introduce a claim-specifi...

A. Alzahrani · 0 citations
#artificial intelligence Preprint Sep 2026

Who Audits Whom, on What Substrate, with What Evidence? An Independence-Graded Audit Protocol for Agentic AI

The model is given a formal basis by transplanting the beta-factor model of common-cause failure from reliability engineering, a seven-step protocol whose outputs a third party can verify, a structural detectability analysis of a procurement-controls agent audited at three grades, and a Monte Carlo study of the model.

Mohamed Chahine Ghanem · 0 citations
Preprint Sep 2026

Audit Without Verification: When LLM Accountability Layers Relay Rather Than Check

Multi-agent LLM pipelines increasingly span organisational boundaries; when a fault surfaces, someone must determine where it entered. The artifact available is rarely a full execution trace: it is the reports each agent filed, and a filed report can state a conclusion alongside its observations. Using a pre-registered...

Paul-Peter Arslan · 0 citations
Preprint Aug 2026

Governing Agentic AI in FinTech

This work develops a multilevel governance theory for agentic AI and test its mechanisms in three studies over nine model versions, from a three-billion-parameter local model to a commercial frontier system.

Henry L. Han · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.