Skip to content
Preprint

Agent Flight Recorder: Tamper-Evident Audit Trails with On-Chain Anchoring for Long-Horizon Tool-Using Agents

Sep 2026 · 4 citations · 31 references
Computer Science

TL;DR

The Agent Flight Recorder captures each agent action as a structured, canonically serialized event binding eight semantic fields from intent through execution to provenance, and periodic on-chain anchoring of epoch roots lets any verifier with the disclosed payload and Merkle proof check the record independently, without pre- agreeing on a trusted intermediary.

Abstract

Long-horizon agents execute thousands of actions, resulting in sequential failures rather than isolated errors. When a coding agent deletes a production database or a prompt injection spreads across agents, the incident raises questions of causality, authority, and non-repudiable third-party verification. The Agent Flight Recorder captures each agent action as a structured, canonically serialized event binding eight semantic fields from intent through execution to provenance. Hash chaining and Merkle batching provide tamper evidence and compact inclusion proofs. For cross-organizational disputes where no party's infrastructure qualifies as neutral ground, periodic on-chain anchoring of epoch roots lets any verifier with the disclosed payload and Merkle proof check the record independently, without pre-agreeing on a trusted intermediary. The on-chain footprint is minimal: each anchor stores a 32-byte epoch root and a back-pointer, and no event content touches the chain. We evaluate the system across five cumulative ablation configurations on synthetic agent workloads. The full system adds ~48 microseconds median per-event latency and 512 bytes per event. L2 anchoring costs $2.30 per 100K events at 100-event epochs. The full integrity stack detects edit, delete, reorder, and fork tampering at 100% with zero false positives. Structured forensic queries achieve 1.0 precision on guardrail and delegation lookups where unstructured text search yields 0.013 and 0.077 respectively.

View source

Similar papers

Preprint Aug 2026

The Acknowledgment Point Is the System: Durable Policy-Decision Receipts for AI Audit Evidence

The resulting research prototype binds each deterministic policy decision to the exact policy source, commits a privacy-minimizing record at a caller-selected synchronization boundary, and returns an Ed25519-signed receipt that states whether that boundary completed.

Neeraj Kumar Singh Beshane · 0 citations
#artificial intelligence Preprint Sep 2026

Janus: Evidence-Before-Effect Sagas and Offline-Verifiable Provenance for Agentic LLMs

Agentic large language models (LLMs) now move money through tools, yet the record of what they did is usually a trace their own process emits beside the effect. Janus puts the record on the effect path. A step's proposal, the verdict on it and any answer from a validator or a person are durable in a signed, hash-chaine...

Mustafa Arslan · 0 citations
#artificial intelligence Preprint Sep 2026

VST: Verifiable Structured Transport for Auditable Agent-to-Agent Alpha Discovery

Single-run agent-to-agent alpha discovery results are reported descriptively, gross of costs, and are explicit about their limits throughout; in particular they do not isolate the effect of the leap machinery from the inherited search substrate, which is left to future work.

Yu-Qi Li, Siyuan Liu, Bing-Jun Liu · 0 citations
#artificial intelligence Preprint Sep 2026

Actions with Receipts: Jointly Binding Claims, Evidence, and Execution for Replayable Tool-Agent Auditing

Tool-using agents can expose citations and execution logs while leaving a critical association unaudited: whether the claim shown to a user is the claim emitted by the committed execution and supported by the cited source. A valid citation and a valid trace can therefore remain individually well formed while being tran...

Miao-Bo Hu, Shu-Hao Hu, Xiao-Bo Guo et al. · 0 citations
Preprint Aug 2026

Reconcile Once, Write Anytime: A Trust-Tiered Librarian and a Multi-Agent Writer for Drift-Free, Point-in-Time Research

Long-form research reports generated by large language models drift, contradict themselves, and lose provenance: the same metric appears with different values, and rumor is quoted as confidently as an audited filing. We present a two-tier agentic system that separates a maintained, point-in-time knowledge library from...

Xing Zhang, Ya Cui, Guanghui Wang et al. · 0 citations
Conference Open access Sep 2026

When Event-Driven Systems Become the Source of Truth: Lessons from CDC, Kafka, and Replay-Safe Inventory Platforms

This paper examines the point at which event streams stop being helpful integration plumbing and become operational truth for business-critical state. Using CDC, Kafka and replay-safe inventory platforms as the frame, it focuses on the failures that matter after the brokers are healthy: ambiguous replay intervals, dupl...

Ishan Shah · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.