Skip to content

EvolveTrade: Experience-Driven Policy Refinement for Self-Evolving LLM Trading Agents

Sep 2026 · 0 citations · 38 references
Computer Science

TL;DR

Experiments show that EvolveTrade often improves Sharpe Ratio and Cumulative Return over fixed-policy LLM baselines, achieving the improved SR and CR in most evaluated settings, and suggest that adapting the reusable procedure governing tool use is a key direction for building more robust LLM trading agents.

Abstract

Large language model (LLM) trading agents can combine market data, news, and executable analysis, but their behavior is often controlled by static hand-written tool-use policies that are fixed before deployment. This limits their ability to adapt how they gather evidence, invoke tools, verify signals, and manage risk under changing market regimes. We introduce EvolveTrade, a self-evolving framework that treats the system prompt of a tool-using trading agent as a text-parameterized policy. After each update interval, a Policy Agent revises this policy using accumulated decision traces and realized portfolio feedback, while keeping the backbone LLM fixed. The updated policy is then used for the next batch of trading decisions, enabling the agent to refine its information-acquisition and portfolio-construction procedure over time. Experiments across multiple market regimes and two LLM backbones show that EvolveTrade often improves Sharpe Ratio and Cumulative Return over fixed-policy LLM baselines, achieving the improved SR and CR in most evaluated settings. Behavioral analyses further show that self-evolved policies increase code-mediated analysis and activate regime-relevant computations; case-level policy-to-return attributions trace how policy-induced allocation changes contribute to realized return differences. These results suggest that adapting the reusable procedure governing tool use is a key direction for building more robust LLM trading agents.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

LiveOption: Evaluating LLM Agents in Structured Option Trading with Nonlinear Payoffs

Large language models (LLMs) and multi-agent systems (MAS) have shown promise in financial decision-making, yet existing evaluations focus on equity trading and primarily assess directional prediction, overlooking the structural complexity of derivative markets. Option trading introduces fundamentally different challen...

Hao-Chen Luo, Yi-Fan Li, Binh Minh An et al. · 0 citations
Review Open access Sep 2026

GroundTrader: Multi-Agent Transition Supervision for LLM-Based Financial Trading

Large language model (LLM)-based financial agents use market information to generate trading actions together with textual reasoning. However, the resulting executable position transition may conflict with its supporting analysis, while the effect of intervening on that transition becomes observable only after the mark...

Xin-Dai Cui, Yu-Tong Liu, Shu-Lei Zhang · 0 citations
Sep 2026

When AI Meets Finance (StockAgent): A Benchmark for Simulating Large Language Model Behaviors in Controlled Trading Environments

Can AI Agents be benchmarked within strictly controlled simulated trading environments to investigate how external factors impact their collective trading behaviors? These factors, which frequently influence trading behavior, are critical elements in the quest to maximize investors’ profits. Our work aims to address th...

Chong Zhang, Xinyi Liu, Zhongmou Zhang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

What LLM Trading Agents Actually Do in Production: A Six-Month, Population-Scale Record from Two Fleets

We present a continuous, population-scale measurement record of autonomous language-model trading agents operating in production across two systems with one design lineage: DX Terminal Pro (3,505 user-funded vaults trading real ETH in Base memecoin markets for 21 days, February to March 2026) and the DXAP live alpha fl...

T. Barton, Chris Constantakis, Patti Hauseman et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Beyond Skill Evolution: Self-Evolving Context Management Policies for Long-Horizon Agent Harnesses

Harness evolution improves LLM agents by learning from execution trajectories, but existing experience- and skill-based methods are less effective on long-horizon tasks. As interactions grow, useful evidence can be buried by redundant or outdated context, making context management itself a key bottleneck. We introduce...

Wei-Yuan Li, Jing-Heng Xu, Ai-Li Chen et al. · 0 citations
Preprint Aug 2026

IAPO: Influence-Aware Policy Optimization for Credit Assignment in Multi-Turn Service Agents

Influence-Aware Policy Optimization (IAPO), which represents each rollout as a typed influence-dependency graph over trainable agent actions, with user and tool observations serving as evidence, is introduced and advances the understanding of credit assignment in multi-turn user interactions.

B. Ren, Yirong Mao, Yi Yang et al. · 1 citation

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.