Skip to content
Preprint

The Working Set of a Coding Agent: Coherence Debt in Repository-Scale Tasks

Aug 2026 · 3 citations · 36 references
Computer Science

TL;DR

This work reconstructs a coupled-fact graph as reconstructing a coupled-fact graph: at each edit, a required fact comes from recent context or parametric memory, and the facts covered by neither form coherence debt.

Abstract

Repository-scale coding requires an agent to keep tests, imports, configuration, and migration rules consistent within a bounded context window. We model this as reconstructing a coupled-fact graph: at each edit, a required fact comes from recent context or parametric memory, and the facts covered by neither form coherence debt. We supply and withhold each channel and inject faults across seven models and five harnesses. As expected, no model completes a task on an unseen API with both channels empty, and putting the facts in the prompt restores success. When a rename defeats what models memorized about a real library, all seven fail in the same place, passing and missing the same tests. Availability decides the outcome and distance does not: withholding a fact costs exactly the work it supports, and a supplied fact works as well far from the edit as next to it. Harnesses pay unequal prices for it: configurations that all pass every test differ more than tenfold in tokens consumed because they rebuild the same content at different rates, and spending more recovers nothing when facts are withheld. A missing fact produces wrong work rather than absent work: an agent asked to act acts, fabricating the file or guessing the value, so instruments built on reads look for a hole already filled. How often it says it is blocked instead is a property of the model, from every trial to none. Availability does not settle every edit: where standard and code disagree, agents follow the standard even when it prescribes the worse code, so a stale convention file costs more than no file. Because parametric memory substitutes for reading, on SWE-bench, where models likely know the repositories, reads no longer predict success. Harnesses should keep the facts an edit depends on available when the agent writes, and check that availability against what the agent produces rather than what it reads.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Skill Issue: Lessons from Optimizing Repository SKILLs for Coding Agents

Coding agents increasingly read repository knowledge from SKILLs --- plain \texttt{.md} files versioned alongside the code. Recent work synthesizes these files automatically, by optimizing the document against a benchmark. A bare repository comes with no benchmark, and the synthetic tasks prior work builds are small en...

M. Kozyrev, A. Kozyrev, A. Podkopaev · 0 citations
#artificial intelligence Preprint Sep 2026

StateTape: Action-Conditioned Evidence Lifecycle Modeling for Long-Horizon Coding Agents

Despite the recent success of coding agents built on large language models, it remains challenging to run them over long horizons, since every observation is appended to the context and the context grows with each one. History-based maintenance is a common remedy, which masks or summarizes old observations, or prunes w...

Zi-Yang Yu, Liang Zhao, Bo-Wen Zhu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Closing the Consistency Gap: Self-Evolving Agents That Learn to Stay on Course

Large language model (LLM)-powered agents can be accurate on average yet unreliable in production, a discrepancy that has been observed but remains largely unaddressed. When given the same task five times, a ReAct agent on the AppWorld benchmark using GPT-4.1 succeeds in all five runs only 53% of the time, even though...

Evelyn Duesterwald, Benjamin Elder, Lilian Ngweta et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SkillGLoW: Procedural-Family Skill Consolidation for Self-Improving Agents on Long-Horizon Task Streams

This work argues the missing unit of reuse is the solving procedure shared by a cluster of related tasks, and builds SkillGLoW (Global-Local Weave) around it: the local skills a task writes from its own execution are aggregated into procedural families and compressed into de-instantiated global priors.

Ao Yan, Xin Zhang, Jiawei Du et al. · 0 citations
Preprint Sep 2026

Scanning the Harness: An Empirical Study of Supply-Chain Defects in AI Coding-Agent Configurations

AI coding agents such as Claude Code, Cursor, GitHub Copilot, and OpenAI Codex are configured through artifacts developers write and share: instruction files, skills, hooks, MCP server declarations, subagents. This harness is a dependency layer installed from marketplaces and public repositories, running with the devel...

Benjamin Kapner, Carmel Soceanu, A. Petrunin et al. · 0 citations
#natural language process... Preprint Sep 2026

Compile, Don't Memorize: A Context Compilation Architecture (CCA) for In-Context Learning

The Context Compilation Architecture (CCA), whose central novelty is a typed intermediate representation (IR) with fixed slots (rules.{must_do, must_not, conditional}, output_spec, available_tools, data_profile) into which any prose context is compiled once; executable verifiers and a violation-gated correction loop fo...

Jin-Hu Qi, Min-Da Hu, Wen-Tao Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.