In harness self-evolution, agents modify their own prompts, code, tools, and orchestration while keeping the underlying language model fixed. Recent work has shown that agents can improve themselves in response to task failures and achieve substantial performance gains. However, gains on failed tasks do not automatical...
Qi Cai, Yong-Gang Zhang, Jun Nie et al.· 2 citations· ⚡2
AtmosCoder-Bench is introduced, an execution-grounded benchmark that makes the calculation process visible, and finds that multiple-choice formats inflate measured accuracy by at least 12 percentage points.
Mao-Hao Ran, Chendong Ma, Yanting Zhang et al.· 0 citations
It is proved that the planner's suboptimality is bounded by twice this discrepancy between the predicted and the true plan-cost at the plan the planner commits to, whereas the data-averaged prediction error neither bounds nor tracks it.
Hanzhe You, Yonggang Zhang, Maohao Ran et al.· arXiv.org· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.