Skip to content
Review

Scanning the Harness: Configuration Exposures in AI Coding-Agent Supply Chains

Sep 2026 · 0 citations · 30 references
Computer Science

TL;DR

The study provides reproducible evidence about repository declarations, and identifies concrete opportunities to pin dependencies, review execution pre approval, and check component conformance before distribution or use.

Abstract

AI coding agents rely on repository instructions, skills, hooks, tool-server declarations, and subagent definitions. These artifacts distribute both behavior and access to executable dependencies, making configuration review part of the agent software supply chain. We study 3,171 public GitHub repositories: 2,660 assembled setups and 511 skill collections. Deterministic analysis, mechanical re-derivation, model-assisted adjudication, human review, and platform documentation checks identify six categories of configuration exposure and conformance issues. Unpinned MCP package declarations occur in 9.8% of setups, broad execution grants in 2.5%, and broad skill tool preapproval in 3.8%. Their union covers 409 setups (15.4%); among setups with MCP configuration, 24.5% contain an unpinned declaration. Including required-field and skill-format issues brings the setup rate to 17.9% and the collection rate to 6.8%. Holding the six categories fixed, contextual review changes the setup rate from 18.3% to 17.9%; rule selection explains most of the reduction from the broader candidate set. The findings identify concrete opportunities to pin dependencies, review execution pre approval, and check component conformance before distribution or use. A documented permission exception also exposes a shared error in the scanner and its mechanical audit, motivating version-specific semantic checks. The study provides reproducible evidence about repository declarations; agreement with human judgments informs label validation, while runtime consequences and recall remain unmeasured.

View source

Similar papers

Preprint Sep 2026

Scanning the Harness: An Empirical Study of Supply-Chain Defects in AI Coding-Agent Configurations

AI coding agents such as Claude Code, Cursor, GitHub Copilot, and OpenAI Codex are configured through artifacts developers write and share: instruction files, skills, hooks, MCP server declarations, subagents. This harness is a dependency layer installed from marketplaces and public repositories, running with the devel...

Benjamin Kapner, Carmel Soceanu, A. Petrunin et al. · 0 citations
Preprint Sep 2026

A Finger on the Scale: Covert Policy Steering through Agentic Skills

SkillShift is presented, a constrained black-box framework for covert policy steering without explicit target command injection or task hijacking that combines semantically plausible policy edits with hierarchical validation, failure-guided optimization, and strategy compression to preserve effectiveness, output validi...

Jia-Rui Li, Jia-Hao Chen, Chun-Yi Zhou et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Who Holds the Pen? Let Specifications, Not Agents, Sign Off

Large language model agents increasingly combine generation, decision-making, execution, and self-evaluation within a single agentic loop. Although they operate under external specifications such as task instructions, guidelines, output schemas, and reusable skills, these specifications typically remain context for the...

Hai-Qing Li, Xin-Yu Ma, Yin-Hao Wu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Incident-Arena: Getting agents to the last nine of reliability

AI coding agents are ubiquitous in engineering workflows amongst industry and academia. Yet, despite their use in app coding, relatively less attention has been paid to their ability to execute on production incident response. This emerging field, termed agentic site-reliability-engineering (SRE) contains benchmarks li...

Andre Fu, Malik Drabla, Leon Liu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

When Agents Implement Systems: A Case Study in Defects, Detection, and Evaluation Rigor

A case study of one LLM coding agent implementing a multi-component data system against a detailed pre-existing specification, and evaluates the one retrieval trade-off specified in that architecture: restricting candidates to a graph-identified entity set before ranking versus unfiltered search.

Phanindra Reddy Madduru · 0 citations
#artificial intelligence Preprint Sep 2026

How Do Agent Harnesses Create Value? Planning Information and Release Control in Stateful LLM Agents

Which component matters more depends on the loss assigned to erroneous acceptance: at low liability the planning gain dominates; at high liability the verifier's avoided false passes dominate; and a standalone verifier captures nearly all the false-pass benefit of the full planning-plus-verification stack at a fraction...

Yu-Kun Zhang, Ke-Mu Xu, Yi-Shen Chen · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.