Skip to content

DREvo: Distilling Recalibrated Historical Experience for Harness Self-Evolution

Jul 2026 · arXiv.org · Vol abs/2607.26722 · 4 citations · 22 references
Computer Science

TL;DR

A new harness self-evolution method, named DREvo, is proposed, which integrates function-level evidence anchoring, state-dependent evidence recalibration, and role-conditioned search intent distillation to determine which historical evidence remains valid and where the harness should evolve next.

Abstract

Harness plays a critical role in large language model agent performance, and building a high-performing harness requires substantial expert effort. Therefore, recent research has increasingly explored harness self-evolution, which iteratively proposes, evaluates, and improves harnesses using historical trial experience. However, accumulated historical experience does not always translate into stable search guidance, and performance often fluctuates substantially across evolution iterations, making it difficult to reliably discover high-performing harnesses under a limited evolution budget. We identify two limitations in how existing harness self-evolution methods leverage historical experience: (1) Lack of dynamic reassessment of whether historical experience remains valid for the current harness, and (2) Lack of explicit mechanisms for translating valid historical experience into actionable search directions. To address these limitations, we propose a new harness self-evolution method, named DREvo, which integrates function-level evidence anchoring, state-dependent evidence recalibration, and role-conditioned search intent distillation to determine which historical evidence remains valid and where the harness should evolve next. Under limited evolution budgets, DREvo exhibits smoother evolution trajectories, achieves the highest accuracy on all five benchmarks, and delivers average gains of 16.2% and 14.2% over the evaluated baselines on domain reasoning and agentic tasks, respectively.

View source

Similar papers

Preprint Sep 2026

A Theory of Reliable Self-Evolution for Agent Harnesses

In harness self-evolution, agents modify their own prompts, code, tools, and orchestration while keeping the underlying language model fixed. Recent work has shown that agents can improve themselves in response to task failures and achieve substantial performance gains. However, gains on failed tasks do not automatical...

Qi Cai, Yong-Gang Zhang, Jun Nie et al. · 2 citations · ⚡2
Preprint Aug 2026

Evo-Bench: Can Language Models Improve Agent Harness?

Evo-Bench is the first benchmark designed to evaluate models'intrinsic harness-evolving capabilities across Search, Office, and General agent domains, and exposes critical temporal anomalies like early saturation, while demonstrating that the synthesized harnesses act as highly transferable reasoning structures, consis...

Lisheng Huang, Chen Yang, Hao Zhou et al. · 6 citations · ⚡1
#natural language process... Preprint Sep 2026

Experience Funnel: A State-Policy Alternating Loop for Self-Evolving Agents

Autonomous agents powered by large language models (LLMs) continuously accumulate experience through interaction, creating an opportunity to improve future behavior through self-evolution. A fundamental challenge is how to transform abundant, task-specific interaction experience into reusable model competence without s...

Wen-Bo Gao, Zhao-Mou Song, Zhi-Yuan Ji et al. · 0 citations
#artificial intelligence Preprint Sep 2026

HarnessEvolve: Learning from Reference Trajectories for Reliable Agent Self-Evolution

This paper introduces HarnessEvolve, a self-evolving framework that learns from reference trajectories to achieve reliable agent self-evolution, and overcomes credit assignment failure by generating reference trajectories and aligning failed executions against them to extract error signals.

Wen Jiang, Ming-Min Chu, Yiding Tian et al. · 5 citations
#artificial intelligence Preprint Sep 2026

Beyond Endpoint Performance: Process-Level Evaluation of Self-Evolving Agents

Self-evolving agents convert interaction feedback into persistent artifacts, such as memories or skills, which in turn guide subsequent decisions. As these artifacts are iteratively updated throughout an experience stream, the capabilities they support may evolve. Consequently, endpoint performance alone offers an inco...

Hong-Qiang Lin, Chao Liu, Xiao-Fan Bai et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.