Skip to content

Trajectory-Aware Benchmark Subset Selection for Cost-Efficient Software Engineering Agent Regression Testing

Sep 2026 · 0 citations · 60 references
Computer Science

TL;DR

This work proposes a trajectory-aware subset selection approach that replaces random sampling with deterministic selection based on trajectory embeddings, and shows that a 10% trajectory-aware subset keeps the median estimation error below 5% while cutting token cost by roughly 90%.

Abstract

Autonomous software engineering agents (SWE-agents) automate coding tasks. Each agent update may require re-running the full benchmark to detect regressions and improvements, at a cost of hundreds of millions of LLM tokens per run, which makes evaluation a bottleneck. One solution is to evaluate only a subset of benchmark instances. Yet, simple approaches, such as random sampling or stratified random sampling based on past pass/fail outcomes, risk producing high variance and unrepresentative subsets. We turn to agent trajectories, the step-by-step record of the actions an agent took. We propose a trajectory-aware subset selection approach that replaces random sampling with deterministic selection based on trajectory embeddings. We first group test set instances by their test outcome in a recent full test run to preserve the historical pass/fail rate, then select the subset using the trajectory's embedding space. We evaluate 76 subset selection configurations, including random sampling, embedding-based selection, clustering-based selection, and hybrid shortlist-then-subsample strategies, across three regression scenarios: same-configuration reruns, model and configuration changes, and agent framework changes. Our best trajectory-aware method is the one selecting benchmark instances closest to the centroid of each outcome group in the embedding space. It achieves the lowest estimation error among all methods we evaluate. For instance, when evaluating a given agent version on a selected subset of 5% or 10% of the test instances, our approach reduces the average estimation error by 3--11% and the worst-case error by 4--11% relative to the typical draw and 38--46% relative to the 95th-percentile draw of the strongest baseline. Our results show that a 10% trajectory-aware subset keeps the median estimation error below 5% while cutting token cost by roughly 90%.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Efficient SWE Agent Benchmarking via Trajectory-Aware Evaluation

PTA-IRT is proposed, a Privileged Trajectory-Aware Item Response Theory framework that fuses process and outcome signals and consistently outperforms prior IRT baselines on score and ranking recovery across four SWE benchmarks.

Ke-Feng Duan, De-Wu Zheng, Yan-Lin Wang et al. · 0 citations
#software testing Preprint Aug 2026

Disagree to Explore, Agree to Commit: Routing-Guided Test-Time Scaling for Software Agents

This work introduces Risa (Routing-Informed Steering and Arbitration): within trajectories, routing encourages diverse exploration and controlled convergence during patch commitment; across separately sampled trajectories, agreement at informative patch positions selects a final candidate.

Kang Chen, Junjie Nian, Yi-Xin Cao et al. · 0 citations
#artificial intelligence Preprint Sep 2026

TraceDance: An Automated System for Building Agent Behavior Benchmarks from Real-World Agent Deployment Traces

An agent can complete a task while exhibiting undesirable behavior during execution. Developers need tests for the specific behaviors encountered in deployment, beyond fixed benchmark suites. We present TraceDance, an agent system that constructs targeted benchmarks from deployment traces for user-specified undesirable...

De-Hai Min, Dao-An Zhang, Yiming Zeng et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Substrate-Aware AI Agents: Execution Context as a First-Class Input

A minimal execution contract induces proactive structural adaptation in generated programs, shifting computation away from unconstrained allocations and substantially improving observed resource-time profiles before execution, establishing a controlled proof of concept for substrate-aware agent planning.

Manu Agrawal · 0 citations
#artificial intelligence Review Sep 2026

RankEvolve: A Reliable Multi-Agent Auto-Research Harness for Evolving Ranking Models

Auto-research agents, LLM systems that propose, implement, train, and evaluate model changes across iterations, promise to automate applied ML's experimental loop. Over long horizons, execution accuracy is a binding constraint: a change can silently leak held-out data, omit normalization, disconnect a gradient, or leav...

Zheng-Yu Chen, Lin-Feng Liu, Hong Li et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Oct 2, 2026

Documenting the tech worker movement

Writing as a participant and researcher, PhD student JS Tan SM ’22 has co-authored a new book about the rise of tech worker protests and the employer backlash that followed.

GPT-Lab Sep 23, 2026

Requirements Don’t Live in Isolation: What We’re Exploring with Req-Space

Requirements in large systems rarely exist in isolation. Their meaning depends on the wider project context - other requirements, policies, decisions, tests, and implementation details. That becomes especially important when AI is used for review, because spotting a possible conflict or gap is only the beginning. ReqSpace explores how AI, visualisation, and connected project context can help reviewers understand those findings, trace the relationships behind them, and focus on the questions that…

GPT-Lab Sep 17, 2026

Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering

AI is making software generation faster, but speed does not remove the need for expertise. As more work is delegated to AI, tacit knowledge may become one of the most important human advantages in software engineering. The post Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering appeared first on GPT-Lab.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.