Skip to content

Self-Designed Evaluators and Warm Memory for Long-Horizon Agents

Sep 2026 · 0 citations · 101 references
Computer Science

TL;DR

SelfSuite is presented, in which the agent's own base model, given only the world's public materials, designs a small evaluation suite of weighted judges and grounded per-task briefs, freezes it, and uses it to gate a keep-best retry and to label a typed, outcome-tracked memory.

Abstract

A tool-using language-model agent deployed over a long stream of tasks receives no reward, so it cannot tell whether it succeeded, cannot safely retry, and cannot label the experience it needs to improve. We present SelfSuite, in which the agent's own base model, given only the world's public materials, designs a small evaluation suite of weighted judges and grounded per-task briefs, freezes it, and uses it to gate a keep-best retry and to label a typed, outcome-tracked memory. On matched five-repeat benchmarks over tau2-bench and AppWorld, SelfSuite scores above the plain agent without any labels, matches methods given ten expert labels on tau2-bench, and trails Agentic Context Engineering (ACE) on AppWorld, where code execution gives a direct success signal. In an ablation campaign run on the same tasks, it is above label-free ACE in every repeat, and the gated second attempt is the only component whose removal hurts in every repeat. We also simulate a subject-matter expert who grades ten onboarding tasks per world. Using those labels to calibrate SelfSuite's evaluator gives a small, consistent gain, and using them to warm up ACE's memory lifts ACE to tie calibrated SelfSuite. A single-run study on a second model family shows the same ordering.

View source

Similar papers

#machine learning Preprint Oct 2026

ASCENT: Online Test-Time Training of Long-Horizon Agents via Self-Distillation of Verified Experience

A large language model (LLM) agent solves long-horizon tasks through many reasoning-action turns, with one verification signal at termination. Deployed agents face streams of related tasks, making their trajectories a natural resource for improvement. In-context adaptation agents store reflections, memories, or skills...

Hao-Dong Lu, Dong Gong · 0 citations
#artificial intelligence Preprint Sep 2026

DolphinBench: Mapping the Pareto Frontier of Agent Memory

DolphinBench is presented, a benchmark that evaluates memory directly through an agent's task completion and requires all evaluations to report total cost and latency alongside accuracy, which enables us to evaluate agent memory systems holistically.

Soumil Rathi, Deshraj Yadav, Taranjeet Singh · 0 citations
Preprint Aug 2026

Prime Agent: A Self-Improving RLM Harness

Low-friction, expressive membrane prevents harness failures from becoming model failures and pushes measurement toward the model's true maximal underlying capability, on Factorio, where refinement allows for continuous technology progression and dedicated subagents enable parallelized work.

Seth Karten, Alex L. Zhang, Kevin Thomas et al. · 13 citations · ⚡2
#artificial intelligence Preprint Sep 2026

VRL-Bench: Benchmarking agents on computer control tasks under finite trial budgets

Learning from trial and error is a promising way to improve language agents on complex tasks such as computer control. Reflexion introduced verbal reinforcement learning, which turns failed trials into text that guides later attempts without updating model parameters. We introduce VRL-Bench, a harness for fair evaluati...

Yu Bai, Yukai Miao, Dawei Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SAGE: Symbolic Action-Gating and Editing for LLM Task Planners

SAGE (Symbolic Action-Gating and Editing), a single-LLM planner built from two lightweight mechanisms: a domain-agnostic symbolic gate that blocks precondition-violating actions with typed reasons as a runtime safety monitor, and a local edit that regenerates only the failed sub-goal's suffix, keeping completed and unt...

T. Bui, Jongsul Moon, Youngouk Kim et al. · 0 citations
#artificial intelligence Preprint Sep 2026

How Strongly Should Task State Influence an LLM Agent?

Across three models, two reasoning regimes, and two domains, four findings hold without per-turn reasoning: displaying accurate state is unreliable, an unverified ledger the agent writes itself beats an accurate checklist it is shown, directives help in proportion to the model's obedience, and enforcement needs no obed...

C. Zhang, Wonbin Kweon, Jiawei Han · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.