Skip to content
Book Open access

Large Language Model (LLM) as an Excellent Reinforcement Learning Researcher in both Single-Agent and Multi-Agent Scenarios

Aug 2026 · Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 · 0 citations · 55 references

TL;DR

A Self-Evolutional single-agent/multi-agent Reinforcement Learning (SE-RL) framework that utilizes a Large Language Model (LLM) to design various RL algorithm modules, such as agent model design, reward function, profiling, communication, and state imagination, by leveraging the LLM generating module output or code.

Abstract

In the quantitative finance area, particularly in order execution, reinforcement learning (RL) has shown great promise due to its ability to interact with market environments based on real data. However, traditional RL methods suffer from slow research speed and rely on static market assumptions, which do not consider the impact of the agent's execution action on the environment. To address these, we propose a Self-Evolutional single-agent/multi-agent Reinforcement Learning (SE-RL) framework. The framework utilizes a Large Language Model (LLM) to design various RL algorithm modules, such as agent model design, reward function, profiling, communication, and state imagination, by leveraging the LLM generating module output or code. SE-RL could continuously improve the accuracy of LLM-generated RL algorithms through a dual-enhancement kit at both high-level (prompt refinement) and low-level (parameter fine-tuning). Additionally, we use a multi-agent system to simulate dynamic financial markets, accounting for the impact of order executions on market dynamics. To further enhance training in such a dynamic market, we develop a hybrid environment training method that could rebalance each environment's loss weight. Comprehensive experiments on 200 realistic stock datasets demonstrate that our proposed framework outperforms current state-of-the-art baselines. Project page: https://kdd2026-se-rl.github.io/.

Read PDF

Similar papers

#large language models Review Open access Sep 2026

A Survey on Reinforcement Learning Optimization Methods for Multi-Agent Collaboration of Large Language Models

The integration of reinforcement learning (RL) into the optimization of multi-agent collaboration for Large Language Models (LLMs) is an important combination of two advanced areas, Multi-Agent Systems (MAS) and LLMs. This paper thoroughly examines the main approaches, evaluation standards, recent progress and existing...

Qian-Ling Zhang · 0 citations
Conference Open access 2026

Towards Automated Reinforcement Learning: Applications and Prospects of Large Language Models as Cognitive Components in the Full RL Lifecycle

: Deep reinforcement learning (DRL) has made significant progress in recent years. However, it still faces multiple bottlenecks in real-world applications. Agents suffer from low sample efficiency, and designing effective reward functions remains very difficult. Furthermore, traditional DRL lacks general cognitive abil...

Sheng-Hao Yuan · 0 citations
#machine learning Preprint Aug 2026

Learning Generalizable Behaviors for Terminal Agents

River, a simple training recipe that improves reward quality by filtering low-quality environments and augmenting outcome rewards with process-level behavior regularization is proposed, which achieves the best performance among evaluated open-source RL-trained 8B models across four terminal-agent benchmarks.

Yi-Fan Yao, Bo Pang, Xuan-Phi Nguyen et al. · 1 citation
#reinforcement learning Conference Open access Sep 2026

Progress and Analysis of Optimization for Large Language Models Based on Reinforcement Learning

Large language models (LLMs) are one of the current research focuses in society and have extensive applications in various fields. However, when facing complex tasks such as mathematical reasoning at present, it will exhibit problems such as weak generalization ability. Reinforcement Learning (RL) can effectively optim...

Feng-Rui Tian · 0 citations
Preprint Aug 2026

IB-RL: Isolated Bilateral Reinforcement Learning for Strategic Dialogue Agents

Isolated Bilateral Reinforcement Learning (IB-RL), in which the two roles coevolve through joint rollouts while each role optimizes its own reward through fully independent advantages, action masks, and update paths, produces policies that generalize more effectively to unseen counterparts.

Senhao Wang, Chenghao Cai, Hai-Tao Hu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.