The study provides reproducible evidence about repository declarations, and identifies concrete opportunities to pin dependencies, review execution pre approval, and check component conformance before distribution or use.
Abstract
AI coding agents rely on repository instructions, skills, hooks, tool-server declarations, and subagent definitions. These artifacts distribute both behavior and access to executable dependencies, making configuration review part of the agent software supply chain. We study 3,171 public GitHub repositories: 2,660 assembled setups and 511 skill collections. Deterministic analysis, mechanical re-derivation, model-assisted adjudication, human review, and platform documentation checks identify six categories of configuration exposure and conformance issues. Unpinned MCP package declarations occur in 9.8% of setups, broad execution grants in 2.5%, and broad skill tool preapproval in 3.8%. Their union covers 409 setups (15.4%); among setups with MCP configuration, 24.5% contain an unpinned declaration. Including required-field and skill-format issues brings the setup rate to 17.9% and the collection rate to 6.8%. Holding the six categories fixed, contextual review changes the setup rate from 18.3% to 17.9%; rule selection explains most of the reduction from the broader candidate set. The findings identify concrete opportunities to pin dependencies, review execution pre approval, and check component conformance before distribution or use. A documented permission exception also exposes a shared error in the scanner and its mechanical audit, motivating version-specific semantic checks. The study provides reproducible evidence about repository declarations; agreement with human judgments informs label validation, while runtime consequences and recall remain unmeasured.
AI coding agents such as Claude Code, Cursor, GitHub Copilot, and OpenAI Codex are configured through artifacts developers write and share: instruction files, skills, hooks, MCP server declarations, subagents. This harness is a dependency layer installed from marketplaces and public repositories, running with the devel...
Benjamin Kapner, Carmel Soceanu, A. Petrunin et al.· 0 citations
SkillShift is presented, a constrained black-box framework for covert policy steering without explicit target command injection or task hijacking that combines semantically plausible policy edits with hierarchical validation, failure-guided optimization, and strategy compression to preserve effectiveness, output validi...
Jia-Rui Li, Jia-Hao Chen, Chun-Yi Zhou et al.· 0 citations
Large language model agents increasingly combine generation, decision-making, execution, and self-evaluation within a single agentic loop. Although they operate under external specifications such as task instructions, guidelines, output schemas, and reusable skills, these specifications typically remain context for the...
Hai-Qing Li, Xin-Yu Ma, Yin-Hao Wu et al.· 0 citations
AI coding agents are ubiquitous in engineering workflows amongst industry and academia. Yet, despite their use in app coding, relatively less attention has been paid to their ability to execute on production incident response. This emerging field, termed agentic site-reliability-engineering (SRE) contains benchmarks li...
Andre Fu, Malik Drabla, Leon Liu et al.· 0 citations
A case study of one LLM coding agent implementing a multi-component data system against a detailed pre-existing specification, and evaluates the one retrieval trade-off specified in that architecture: restricting candidates to a graph-identified entity set before ranking versus unfiltered search.
Which component matters more depends on the loss assigned to erroneous acceptance: at low liability the planning gain dominates; at high liability the verifier's avoided false passes dominate; and a standalone verifier captures nearly all the false-pass benefit of the full planning-plus-verification stack at a fraction...
Yu-Kun Zhang, Ke-Mu Xu, Yi-Shen Chen· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.