This paper studies 37,623 provenance-labeled pull requests from five commercial agents and combines the AIDev dataset with 58,792 cached GitHub API responses to measure security smells in added code, structural maintainability, post-merge churn, revert rates, and human review behavior.
Abstract
Autonomous coding agents now open pull requests in public repositories at a scale that was out of reach two years ago, yet little is known about what happens to that code after it lands. This paper studies 37,623 provenance-labeled pull requests (PRs) from five commercial agents (OpenAI Codex, Devin, GitHub Copilot, Cursor, and Claude Code) and a matched human baseline, drawn from 2,807 GitHub repositories between December 2024 and July 2025. We combine the AIDev dataset with 58,792 cached GitHub API responses to measure security smells in added code, structural maintainability, post-merge churn, revert rates, and human review behavior. Three results stand out. First, quality differences are vendor-specific rather than uniform: Codex-authored PRs were reverted about half as often as human PRs (6.1% vs. 11.5%, odds ratio 0.50), while Devin PRs were reverted more often (14.5%, odds ratio 1.31). Second, agent code pooled across vendors was less likely than human code to contain a security smell (odds ratio 0.63), driven by fewer hardcoded credentials and eval-style constructs. Third, review effort concentrates unevenly: Copilot PRs drew the most human reviews and change requests, and Claude Code PRs waited the longest for a first human review (median 12.6 hours). All pipeline code, statistical reports, and figures are released for replication.
AI coding agents now author a large share of pull requests (PRs) merged into popular open-source projects. A merged agent PR is usually considered finished work; yet, prior studies have reported issues in agent code after the merge (e.g., code smells and static-analysis issues). However, little is known about how often...
Wannita Takerngsaksiri, Nhat Duong, Scott Barnett· 0 citations
It is argued that published acceptance figures describe a narrow and unusually demanding slice of GitHub, and it is recommended that studies of agent contributions report the popularity distribution of their repositories and estimate effects within repositories rather than across them.
Perseus Bhavnagri· International Journal for Re...· 0 citations
Coding agents open pull requests (PRs) that claim to speed up software, but studies of human performance fixes say little about how maintainers respond to such a fix or whether its claim holds. From the 71,677 agent PRs of AIDev v4, a text filter and codebook coding by language models and by the authors select 1,262 pe...
Zhen-Yu Qi, Haotang Li, Jin-Fu Chen et al.· 0 citations
Coding-agent evaluation has progressed from resolving individual issues to carrying out long-horizon development, yet task completion is still largely assessed within a single codebase. In software ecosystems, many features and bug fixes require coordinated changes across multiple repositories. We introduce WideSWE to...
Bao-Yi Wang, Xing-Liang Wang, Jin-Yang Wu et al.· 0 citations
Coding agents are increasingly deployed for iterative development on real repositories, yet existing evaluation barely answers a basic question: \emph{do coding agents reuse existing code or reinvent the wheel?} The question matters: every duplicated implementation is a fix applied twice and agents produce code far fas...
Dong-Sheng Ma, Si-Zhe Wang, Xin-Yi Huang et al.· 0 citations
A controlled experiment comparing five informationally equivalent specification formats across six models from three vendor families finds that structured architecture specifications serve as a capability equalizer, with value inversely proportional to model strength and the largest returns for cost-optimized deploymen...
A. Canedo· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.