Skip to content
Preprint

Architecture as Capability Equalizer for Coding Agents

Aug 2026 · 1 citation · 28 references
Computer Science

TL;DR

A controlled experiment comparing five informationally equivalent specification formats across six models from three vendor families finds that structured architecture specifications serve as a capability equalizer, with value inversely proportional to model strength and the largest returns for cost-optimized deployments.

Abstract

LLM-based coding agents generate complete software systems from high-level descriptions, yet little is known about how the format of architecture specifications affects the quality of generated code or whether this effect depends on model capability. We present a controlled experiment comparing five informationally equivalent specification formats (informal prose, Mermaid diagrams with constraints and ADRs, OpenAPI, C4/Structurizr DSL, and TypeScript interface contracts with ArchUnit-style rules) across six models from three vendor families (Anthropic Claude, OpenAI GPT, Google Gemini). Across 90 multi-turn agent trials, specification format shows a strong format x model interaction. On the strongest models (Sonnet 4.6, GPT-5), format barely matters (quality spread 0.17-0.92). On weaker models, format produces spreads of 0.83-2.42 points, with code-proximate formats (OpenAPI, TypeScript contracts) recovering most of the capability gap. Mid-tier models can consume more tokens than frontier models for worse output when they enter compilation debugging loops that stronger models avoid. Self-validation rates collapse from 100% (Sonnet) to 0% (Gemini Flash) across the capability spectrum. TypeScript contracts triple API route coverage for the weakest model (33% to 100%). Structured architecture specifications serve as a capability equalizer, with value inversely proportional to model strength and the largest returns for cost-optimized deployments.

View source

Similar papers

#natural language process... Preprint Sep 2026

LoLBench: Evaluating Coding Agents with Long-Horizon Proposals on Large Software Systems

Modern coding agents can deliver increasingly large repository-level changes, and recent benchmarks reflect this by emphasizing long-horizon tasks with large reference implementations. Many benchmarks evaluate coding agents'implementation capability to produce correct code edits from detailed specifications. However, p...

Yun Peng, Zi-Han Wu, Ze-Yang Zhuang et al. · 0 citations
Preprint Aug 2026

Understanding the Architecture of Coding Agents: An Exploratory Study Using a Research Prototype

This paper presents Ark (Agent Research Kit), a minimal open-source coding agent designed for research and education that preserves the essential architectural mechanisms of modern coding agents while emphasizing simplicity and clarity, and introduces ArkBench, a lightweight benchmark comprising ten representative soft...

Marco Túlio Valente · 0 citations
#artificial intelligence Preprint Sep 2026

Harness or Model? Isolating the Harness Effect in Agentic Coding with a Contamination-Controlled Private Suite

An agentic coding system couples a language model to a harness: the tools, prompts and control flow that turn a chat model into an autonomous software engineer. Vendors ship harnesses tuned to their own models, and practitioners assume the vendor-native pairing solves more tasks. We measure that assumption with paired...

Mohsen Arjmandi · 0 citations
#artificial intelligence Preprint Sep 2026

Who Holds the Pen? Let Specifications, Not Agents, Sign Off

Large language model agents increasingly combine generation, decision-making, execution, and self-evaluation within a single agentic loop. Although they operate under external specifications such as task instructions, guidelines, output schemas, and reusable skills, these specifications typically remain context for the...

Hai-Qing Li, Xin-Yu Ma, Yin-Hao Wu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

An Empirical Study of Harness Design for Coding Agents

This work studies a lightweight coding harness whose execution loop is fixed while three components are varied: planning, action space, and context management, and finds that context management becomes increasingly valuable as the context-window budget tightens.

Run-Ze Fan, Zihao Zhang, Si-Min Ma et al. · 5 citations · ⚡1
Review Sep 2026

Not All Agents Are Equal: Code Quality and Post-Merge Maintenance Across Five Autonomous Coding Agents in the Wild

This paper studies 37,623 provenance-labeled pull requests from five commercial agents and combines the AIDev dataset with 58,792 cached GitHub API responses to measure security smells in added code, structural maintainability, post-merge churn, revert rates, and human review behavior.

Obada Kraishan · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.