Skip to content
Preprint

InfraBench: Evaluating Infrastructure Agents Across Layers, Lifecycle, and Risk

Jul 2026 · 1 citation · 43 references
Computer Science

TL;DR

InfraBench is presented, a benchmark suite for evaluating AI agents on realistic infrastructure tasks across the full system stack and full operational lifecycle with fine-grained risk assessment and shows that even the strongest agent cannot secure a full score across all tasks.

Abstract

Managing modern computing infrastructure has become a steadily harder problem due to the ever-increasing complexity. Recent advances in AI agents create a timely opportunity to automate infrastructure management tasks, but it remains unclear how well such agents can handle real-world infrastructure complexity. We present InfraBench, a benchmark suite for evaluating AI agents on realistic infrastructure tasks across the full system stack and full operational lifecycle with fine-grained risk assessment. Experiments with 15 agent-model configurations show that even the strongest agent cannot secure a full score across all tasks. Mean effective scores range from roughly 40% to 88% (with per-configuration standard errors of 6-12 points), repeating every task three times reveals that top configurations still pass only a fraction of their attempts, and per-check scoring exposes a general failure pattern: agents may routinely satisfy short-term objectives while leaving non-durable changes, broken distributed invariants, unsafe side effects, and uncleaned state behind. INFRABENCH, including its live leaderboard, tasks, and evaluation harness, is publicly available at infraben.ch.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Incident-Arena: Getting agents to the last nine of reliability

AI coding agents are ubiquitous in engineering workflows amongst industry and academia. Yet, despite their use in app coding, relatively less attention has been paid to their ability to execute on production incident response. This emerging field, termed agentic site-reliability-engineering (SRE) contains benchmarks li...

Andre Fu, Malik Drabla, Leon Liu et al. · 0 citations
#artificial intelligence Preprint Oct 2026

SoK: Decentralized Agent Economic Infrastructure

Decentralized agent economies increasingly build a single task from protocols that were designed and secured separately. This creates a simple problem: a workflow can look correct at each step and still produce the wrong outcome. For example, a correct escrow may release payment on an authorized approval that provides...

Rui Sun, Xi-Han Xiong, Qin Wang et al. · 0 citations
Review Open access Aug 2026

A Systematic Survey of Agentic Skills: Architecture, Lifecycle, and Security

Autonomous large language model (LLM) agents increasingly face reliability, context consumption, and execution stability bottlenecks when deployed on complex, long-horizon tasks. While monolithic prompt engineering and stateless tool-calling paradigms struggle to scale, the field is rapidly converging toward \emph{agen...

Sanket Badhe, D. Shah, Priyanka Tiwari et al. · 0 citations
#artificial intelligence Preprint Aug 2026

ComponentBench: Diagnosing Component-Level Failures in Computer-Use Agents

This work presents ComponentBench, a benchmark and diagnostic pipeline for component-level evaluation of computer-use agents on modern web UIs, and introduces a scalable pipeline for auditing realized structural difficulty after implementation and synthesizing structured failure analyses across tasks and component fami...

Tian-Chen Guan, Xinlei Lin, Royce Cheng-Yue et al. · 0 citations
Review Sep 2026

WideSWE: Can Coding Agents Coordinate Changes Across Repositories?

Coding-agent evaluation has progressed from resolving individual issues to carrying out long-horizon development, yet task completion is still largely assessed within a single codebase. In software ecosystems, many features and bug fixes require coordinated changes across multiple repositories. We introduce WideSWE to...

Bao-Yi Wang, Xing-Liang Wang, Jin-Yang Wu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.