Skip to content
Preprint

From Experiments to Decisions: Reusing Evidence in Autonomous Coding Research

Sep 2026 · 0 citations · 15 references
Computer Science

Abstract

Autonomous coding agents can remember an experiment yet carry forward a conclusion it does not justify. We reconstruct how evidence is reused in a 400-task NeuroGolf campaign, with selected wellbore-prediction records from the same operator as cross-domain comparisons. A numerical counterexample exposes an overbroad exclusion; other episodes distinguish what failed, incomplete and revised programs justify doing next. The wellbore records reveal an additional weakness: a coordinator correctly acknowledges a novelty-only rejection, then summarizes it as measured closure. Separately, a prefix-based acceptance gate at three thresholds yields worse target scores than ungated adaptation in the retained experiment. These cases motivate an inspectable handoff linking the tested proposition, its scope, candidate and evaluator identity, check statuses, and reopening condition. A dispatch tree separates investigation, repair, reopening and stopping; a promotion predicate requires that each check both passes and has evidence supporting its use for the requested decision. The practical lesson is to preserve not just experimental results, but their limits on subsequent action.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.