Sep 2026· International Journal for Research in Applied Science and Engineering Technology· 0 citations
TL;DR
It is argued that published acceptance figures describe a narrow and unusually demanding slice of GitHub, and it is recommended that studies of agent contributions report the popularity distribution of their repositories and estimate effects within repositories rather than across them.
Abstract
Autonomous coding agents now open pull requests on GitHub without a person writing the code, and the share of
those pull requests that maintainers merge has become the usual shorthand for how well the agents work. This paper argues that
the shorthand measures the projects as much as it measures the agents. Using the public AIDev dataset, a snapshot of 2,743,854
pull requests written by six agents across 326,798 repositories between December 2024 and October 2025, we find that the merge
rate falls from 91.58 percent in repositories with no stars to 60.05 percent in repositories with a thousand or more. The decline
appears inside every agent separately, and in the five agents with enough volume in the top band to measure it the fall is between
20.3 and 27.6 percentage points, so it cannot be an artefact of which agent is used where. The enriched portion of the dataset
that previously published analyses rely on contains only repositories with at least 100 stars, which we verify directly; it holds 2.7
percent of the decided pull requests and merges them at 73.67 percent against 90.63 percent elsewhere. Holding the repository
fixed with a Mantel-Haenszel estimator changes the picture between agents: of fifteen comparable pairs, two reverse direction
and one loses more than half its effect.
Time to a decision separates the same way, a median of 51 seconds in unstarred repositories against 5.4 hours in the most
popular ones, with a medium effect size. Even within the enriched subset, 75.8 percent of agent pull requests carry no recorded
human review. We conclude that published acceptance figures describe a narrow and unusually demanding slice of GitHub, and
we recommend that studies of agent contributions report the popularity distribution of their repositories and estimate effects
within repositories rather than across them.
This paper studies 37,623 provenance-labeled pull requests from five commercial agents and combines the AIDev dataset with 58,792 cached GitHub API responses to measure security smells in added code, structural maintainability, post-merge churn, revert rates, and human review behavior.
Coding agents open pull requests (PRs) that claim to speed up software, but studies of human performance fixes say little about how maintainers respond to such a fix or whether its claim holds. From the 71,677 agent PRs of AIDev v4, a text filter and codebook coding by language models and by the authors select 1,262 pe...
Zhen-Yu Qi, Haotang Li, Jin-Fu Chen et al.· 0 citations
GitSkills is presented, a dataset of 3,797,117 $\mathrm{SKILL.md}$ files collected from 282,200 public repositories in July 2026, which retains every file occurrence with its repository, path, and content hash.
Giuseppe Destefanis, Daniel Graziotin, Matteo Vaccargiu et al.· 0 citations
AI coding agents now author a large share of pull requests (PRs) merged into popular open-source projects. A merged agent PR is usually considered finished work; yet, prior studies have reported issues in agent code after the merge (e.g., code smells and static-analysis issues). However, little is known about how often...
Wannita Takerngsaksiri, Nhat Duong, Scott Barnett· 0 citations
AI-research agents, or autoresearch systems, combine language models with tools, search, evaluation, and iterative modification of research artifacts. Their public software ecology is hard to compare because repositories, papers, benchmarks, libraries, and companion artifacts are often counted as one population. We con...
Three minimal copying models, one per decision and with a single free parameter each, reproduce the heavy-tailed distribution of how many agents met on a page, the frequency of the pieces from which the agents built their names, and the patchwork of pages that are internally consistent and different from one another.
G. De Marzo, Nicola Alboré, David García· 3 citations· ⚡1
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.