Skip to content
Preprint

Update from Hell: Can Coding Agents Survive Hidden Breakage in Dependency Upgrades?

Aug 2026 · 0 citations · 25 references
Computer Science

TL;DR

Evaluating mainstream coding agents on DEPBENCH, a benchmark consisting of 203 real-world dependency-upgrade tasks across five package ecosystems spanning five language communities, each involving hidden code-level changes that require source code adaptation.

Abstract

Modern software systems rely heavily on third-party dependencies, but upgrading those dependencies remains a costly maintenance activity. Dependency upgrades do not always preserve the function signatures, type systems, APIs, or runtime semantics assumed by existing code. Consequently, developers often need to perform source code adaptations to accommodate dependency-induced changes. However, such code-level changes are often not explicitly communicated to project maintainers, posing a significant challenge to software reliability. Meanwhile, coding agents have emerged as a new form of software development tool and are increasingly adopted by developers due to their automation capabilities. In this paper, we introduce DEPBENCH, a benchmark consisting of 203 real-world dependency-upgrade tasks across five package ecosystems spanning five language communities, each involving hidden code-level changes that require source code adaptation. We evaluate mainstream coding agents on DEPBENCH. The best completed configuration solves only 104/203 tasks (51.2%), with substantial variation across agent harnesses, models, and ecosystems, highlighting an important gap between current agent capabilities and real-world software maintenance needs.

View source

Similar papers

Review Sep 2026

WideSWE: Can Coding Agents Coordinate Changes Across Repositories?

Coding-agent evaluation has progressed from resolving individual issues to carrying out long-horizon development, yet task completion is still largely assessed within a single codebase. In software ecosystems, many features and bug fixes require coordinated changes across multiple repositories. We introduce WideSWE to...

Bao-Yi Wang, Xing-Liang Wang, Jin-Yang Wu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

From Dead Code and Static Requirements to Working Engines: Software Revival with Coding Agents

Can coding agents restore software that no longer runs while preserving its underlying methods, and reconstruct industrial software engines from open specifications? Here we introduce ReviveBench, a benchmark with two task families evaluated by hidden verifiers calibrated against native execution environments, establis...

Tian-Yu Liu, Ding-Yuan Dai, Yu-Fan Du et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Do Coding Agents Reuse Existing Code or Reinvent the Wheel?

Coding agents are increasingly deployed for iterative development on real repositories, yet existing evaluation barely answers a basic question: \emph{do coding agents reuse existing code or reinvent the wheel?} The question matters: every duplicated implementation is a fix applied twice and agents produce code far fas...

Dong-Sheng Ma, Si-Zhe Wang, Xin-Yi Huang et al. · 0 citations
Preprint Aug 2026

Evaluating Agentic Code Repair Capabilities in Distributed Systems

DDBench is introduced, a code-repair benchmark of 60 historical bugs mined from 13 open-source distributed systems, partitioned into three difficulty tiers, isolating the effect of debugging context from model capability.

Yi-Bo Yan, Huijuan Wang, Jun-Zhou He et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Zero2Repo: Can Coding Agents Build Repositories from Scratch?

Coding agents are increasingly asked to build software rather than patch it, yet benchmarks for from-scratch repository construction are mostly limited to a single language and depend on manually curated tasks. We introduce Zero2Repo, a benchmark in which an agent receives a product requirements document, an interface...

Pei Yang, Tian-Yu Shi, Yu-Hang Yao et al. · 0 citations
#software testing Preprint Aug 2026

SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?

SWE Refactor Bench is introduced, a benchmark comprising 20 whole-repository migrations, covering 4 kinds of technical debt, and SWE Refactor Bench is positioned as a rigorous testbed for developing coding agents for reliable whole-repository migrations.

De-Yao Hong, Yi-Zhe Chi, Wen-Yi Li et al. · 3 citations · ⚡1

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.