Skip to content

Can AI Agents Detect and Repair Artifact Drift in Network Experiments?

Sep 2026 · 0 citations · 26 references
Computer Science Engineering

TL;DR

NetArtifactBench is introduced, which tests whether AI agents can repair inconsistent records derived from public network-system artifacts while preserving claims that remain supported, and argues that artifact integrity should become a first-class design and evaluation requirement for AI agents operating on network systems.

Abstract

In recent years, AI agents have evolved into capable assistants that carry out multi-step tasks in digital environments. The network systems community is beginning to explore these capabilities in operational and experimental settings. However, an agent operating in network systems should not be judged solely by whether it completes the immediate task. The experiment record it modifies must also remain trustworthy. We call this property artifact integrity: the record's claims must remain supported by the available evidence, confined to the scope established by that evidence, and traceable through the artifacts that encode their support. To make this property measurable, we introduce NetArtifactBench, which tests whether AI agents can repair inconsistent records derived from public network-system artifacts while preserving claims that remain supported. The benchmark contains 52 instances with injected inconsistencies ranging from direct contradictions to unstated relations spread across several artifacts. We evaluate 23 agent configurations across three general-purpose AI agent runtimes using deterministic scoring. The average contract pass rate is 65.3 % across 5,980 outputs, but no agent runtime exceeds 30 % when repair requires recovering implicit relations and propagating changes across artifacts. These results reveal a sharp boundary between local correction and complete record-level repair. Therefore, we argue that artifact integrity should become a first-class design and evaluation requirement for AI agents operating on network systems.

View source

Similar papers

Preprint Aug 2026

Multi-Agent AI Safety as an Institutional Design Problem

This is the first paper from POLIS, an ongoing research programme studying algorithmic institutions for multi-agent systems, and asks which parts of an AI institution produce safety and how they do it.

X. Abdullah · 1 citation
#artificial intelligence Preprint Sep 2026

Can AI Agents Deliver Verifiable Network-Wide Outcomes Across Authority Boundaries?

AI agents are increasingly involved in network automation, where they can initiate configuration changes through mediated operational interfaces and assess the resulting state. Nonetheless, operational networks usually span many devices and administrative domains. Realizing an operator's intent requires coordinating ag...

Tianzhu Zhang, Chih-Kai Huang, Mei-Kang Qiu · 0 citations
#artificial intelligence Preprint Sep 2026

Closing the Consistency Gap: Self-Evolving Agents That Learn to Stay on Course

Large language model (LLM)-powered agents can be accurate on average yet unreliable in production, a discrepancy that has been observed but remains largely unaddressed. When given the same task five times, a ReAct agent on the AppWorld benchmark using GPT-4.1 succeeds in all five runs only 53% of the time, even though...

Evelyn Duesterwald, Benjamin Elder, Lilian Ngweta et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Can AI Scientists Coordinate at Runtime?

Multi-agent AI scientists have shown improving performance across a diverse range of tasks. Yet a common approach is design-time agentic orchestration, which typically relies on fixed workflows. In contrast, human scientists coordinate and adjust their division of labor at runtime. We therefore ask: can AI scientists a...

Zi-Jian Liu, Yang Luo, Jun-Yu Lu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Explaining AI Agents Through Execution Traces

AI Agents are increasingly deployed in real-world settings, where they interact with external tools and make sequential decisions with limited human oversight. This creates a pressing need for reliable and auditable explanations of what an agent did and why. However, traditional Explainable AI (XAI) methods fall short...

Vittoria Vineis, Fabiano Veglianti, Lorenzo Antonelli et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Recoverability as a System Primitive for Long-Horizon AI Agents

AI agents can be interrupted while editing files, calling tools, or carrying out multi-step tasks. Restarting repeats completed work, but continuing from unverified or outdated progress can carry earlier errors forward. A saved state is not necessarily a suitable place to resume. We introduce recoverability as a system...

Zhi-Hui Zhang, Wei Liu · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.