Skip to content

TimeEvo: Failure-Driven Self-Evolution of a Time Series Agent

Sep 2026 · 0 citations · 46 references
Computer Science

TL;DR

TimeEvo is proposed, which clusters an agent's diagnosed failures into capability gaps, plans a measurement for each, synthesizes evidence-only tools that fill them, and admits the candidate library only through a paired admission gate.

Abstract

Time series agents answer analytical questions by calling external tools, and which tools they carry is decided by people before the agent runs. However, we identify two failures in this setup. Human-Agent Tool Misalignment: a library of 21 expert-curated tools helps on some tasks and hurts on others, dropping anomaly accuracy under every backbone we test. Silent Harm: one round of generic self-revision changes 147 answers and breaks 56 of them, while the final score moves by less than a point. Both follow from the same gap: whether a tool helps is decided question by question at runtime, while tools are supplied in advance and judged by a single average. To address this, we propose TimeEvo, which clusters an agent's diagnosed failures into capability gaps, plans a measurement for each, synthesizes evidence-only tools that fill them, and admits the candidate library only through a paired admission gate. Experiments on ten time series QA tasks and three backbones show that TimeEvo, starting from an empty library, improves accuracy on every task and every backbone, and that a library grown on a cheap model still gains when it is installed into stronger ones. Code is available at https://github.com/Muyiiiii/TimeEvo.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Closing the Consistency Gap: Self-Evolving Agents That Learn to Stay on Course

Large language model (LLM)-powered agents can be accurate on average yet unreliable in production, a discrepancy that has been observed but remains largely unaddressed. When given the same task five times, a ReAct agent on the AppWorld benchmark using GPT-4.1 succeeds in all five runs only 53% of the time, even though...

Evelyn Duesterwald, Benjamin Elder, Lilian Ngweta et al. · 0 citations
Preprint Aug 2026

TimeSage-EV: A Live Benchmark for Agentic Time Series Analysis in Evolving Environments

TimeSage-EV is introduced, a live benchmark for agentic time series analysis in evolving environments that evaluates state identification, data summarization, and outlook reasoning, and is released as a research resource with monthly updates, code, a leaderboard, and failure-mode analyses.

Qingren Yao, Ya-Xuan Kong, Yuqi Nie et al. · 0 citations
#machine learning Preprint Sep 2026

CoArena: Evaluating Computer-Use and Multi-Agent Systems in Real Time

This work defines real-time as five measurable properties, each with an equation and a worked example: continuous task arrival, live concurrent execution, online rating updates, freshness with contamination resistance, and bounded feedback latency from a failed run to a reusable training environment.

Nitish Kovuru, Prateek Jannu · 0 citations
#artificial intelligence Preprint Sep 2026

DolphinBench: Mapping the Pareto Frontier of Agent Memory

DolphinBench is presented, a benchmark that evaluates memory directly through an agent's task completion and requires all evaluations to report total cost and latency alongside accuracy, which enables us to evaluate agent memory systems holistically.

Soumil Rathi, Deshraj Yadav, Taranjeet Singh · 0 citations
#artificial intelligence Preprint Oct 2026

Agents Are Systems, Not Models: Rethinking Agentic Evaluation

Agent evaluations increasingly go beyond a single success rate, reporting metrics such as cost, consistency, and robustness. Yet they typically treat the agent itself as fixed. In practice, an agent is a configurable system: users decide what to tell it, how long to let it run, and which model to use, and each of these...

Luis Wiedmann, Leander Girrbach, Cordelia Schmid et al. · 0 citations
Preprint Aug 2026

EvoTS-Agent: A Self-Evolving LLM Agent for Financial Time Series Change Point Detection

Financial time series exhibit non-stationary and heterogeneous statistical properties, making change-point detection challenging because no single unsupervised algorithm performs consistently across assets and market regimes. Conventional workflows consequently depend heavily on expert-driven model selection, feature d...

Lei Jiang, Yehua Wei, Xinyu Xi et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.