Skip to content

Ecdysis: Efficient and Effective Training of Runtime Harnesses for LLM Agents

Sep 2026 · 1 citation · ⚡ 1 influential · 33 references
Computer Science

TL;DR

This work proposes Ecdysis, which aggregates failure evidence across task instances before promoting recurring failure patterns into persistent harness evolution, biasing evolution toward repairs that are more likely to generalize beyond individual model behaviors.

Abstract

Self-evolving runtime harnesses can substantially improve the capabilities of large language model (LLM) agents and provide a promising paradigm for optimizing agent execution. Existing failure-driven approaches often treat observed agent failures as direct evidence for harness modification. A key challenge in failure-driven harness evolution is that observed failures can reflect either limitations of the underlying model or systematic deficiencies of the harness. Directly optimizing against individual failures can therefore induce model-specific accommodation and impair generalization across tasks and models. We study whether failure evidence accumulated across task instances can provide a more reliable signal for harness training. Our key insight is that failures recurring across distinct tasks provide stronger inductive evidence for systematic harness deficiencies than isolated failures. Based on this insight, we propose Ecdysis, which aggregates failure evidence across task instances before promoting recurring failure patterns into persistent harness evolution, biasing evolution toward repairs that are more likely to generalize beyond individual model behaviors. Ecdysis further employs collaborative failure analysis to refine modification specifications, trading additional evolution-time reasoning for improved modification quality. Across multiple LLMs and benchmarks, Ecdysis improves the reasoning accuracy of evolved harnesses by 18.56% over existing harness evolution while achieving up to 1.84x faster harness training. Ecdysis also enables more data-efficient training. Fine-grained analysis shows that Ecdysis reduces model-specific accommodation during evolution, while the resulting harnesses exhibit stronger cross-LLM generalization and lower inference-time token consumption.

View source

Similar papers

Preprint Aug 2026

Evo-Bench: Can Language Models Improve Agent Harness?

Evo-Bench is the first benchmark designed to evaluate models'intrinsic harness-evolving capabilities across Search, Office, and General agent domains, and exposes critical temporal anomalies like early saturation, while demonstrating that the synthesized harnesses act as highly transferable reasoning structures, consis...

Lisheng Huang, Chen Yang, Hao Zhou et al. · 6 citations · ⚡1
Preprint Aug 2026

AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces

The proposed AutoSaddler, an automatic harness optimization framework that formulates harness improvement as an offline learning problem and iteratively updates the harness using failure signals from mini-batches, suggests that automatic harness optimization is a promising path toward more performant and reliable agent...

Sungho Park, Wonjoong Kim, Rongyuan Tan et al. · 9 citations
Preprint Aug 2026

Verify Smarter, Evolve Further: Efficient Harness Evolution through Behavior-Aware Verification

HarnessLens is introduced, a budget-aware framework for automated harness evolution that jointly explores the task space and user-configurable components, derives candidate modifications from execution trajectories, and selectively verifies each candidate on behavior-relevant tasks using an attributable-evidence gate.

Jing-Heng Xu, Yi-Kai Zhang, Aiden Chen et al. · 7 citations · ⚡1
Review Sep 2026

EvoSafeHarness: Evolving Model- and Domain-Specific Harnesses for Securing Agents

Large Language Model (LLM) agents are turning language into real-world effects, making safety necessary against both indirect prompt injections and direct harmful requests. System-level safety harnesses add an enforcement layer beyond model-level defenses, but existing harnesses are usually designed once by experts and...

Nanxi Li, Ying-Zi Ma, Yulong Cao et al. · 2 citations
#natural language process... Preprint Sep 2026

EvoHarnessBench: Can Your Agents Keep Pace with an Evolving Harness?

Modern LLM-based agents operate through a harness of tools, reusable skills, and specialist agents that shapes what they observe and what they can do. In practice, this harness continually evolves as new capabilities are added. We introduce EvoHarnessBench, a benchmark for evaluating agents under controlled harness evo...

Zi-Xuan Ke, Vaidehi Patil, Hai-Zhou Shi et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Vestrum: Improving Agent Harnesses by Adapting Their Verification, Structure and Memory

An agent harness controls how a language model accesses information, uses tools, preserves memory, and checks its work. Improving this software is costly when each evaluation requires a long interaction with an environment. We introduce Vestrum, a framework that turns failures in execution traces into scoped harness ch...

Jayant Parashar, Eugene F. Douglass, William C. Bastian et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.