Skip to content

AHRL-PM: Asynchronous Hierarchical Reinforcement Learning Framework for Enhanced Portfolio Management.

Jul 2026 · IEEE Transactions on Neural Networks and Learning Systems · Vol PP, pp. 1-14 · 0 citations
Medicine

TL;DR

Ablation studies comparing the AHRL-PM model to a synchronous model and single-layer reinforcement learning (RL) approaches validated its superior performance in terms of profitability and risk-adjusted returns, while also highlighting significant reductions in training time and appropriate portfolio weight adjustments to respond to market dynamics and uncertainties, underscoring the model's efficiency and practical applicability.

Abstract

Effective portfolio management (PM) is a cornerstone of financial strategy, yet it is often challenged by the uncertainties and the high dimensionality of market data. Traditional PM techniques, whether model-based or model-free, frequently fail to address these complexities, resulting in suboptimal asset allocation. This study introduces an asynchronous hierarchical reinforcement learning (AHRL-PM) framework aimed at enhancing active PM. The proposed AHRL-PM approach employs a two-tiered agent system to navigate market complexities. The first-layer agent selects top-performing stocks monthly based on alpha factors, while the second-layer agent adjusts asset weights using market fundamentals and technical indicators, dynamically rebalancing the portfolio in response to market shifts on a daily basis. We rigorously evaluated our AHRL-PM framework through extensive experiments in six globally representative markets. Performance was assessed using two key metrics: annualized return (AR) and Sharpe ratio (SR). Our model consistently outperformed traditional benchmarks on both metrics across all the markets, demonstrating robustness and versatility in PM. Furthermore, ablation studies comparing the AHRL-PM model to a synchronous model and single-layer reinforcement learning (RL) approaches validated its superior performance in terms of profitability and risk-adjusted returns, while also highlighting significant reductions in training time and appropriate portfolio weight adjustments to respond to market dynamics and uncertainties, underscoring the model's efficiency and practical applicability.

View source

Similar papers

Review Open access Jul 2026

Systematic review of reinforcement learning for automated equity portfolio management from single agent to multi agent systems

The systematic development of single-agent to multi-agent ensemble systems shows great improvements in algorithmic and architecture of DRL-based portfolio management, and the research in the future focuses on the importance of explainable AI integration, meta-learning market regime adaptation, and consistent evaluation systems in reproducible research.

Aditi Kumar Rout, U. D. Acharya, Prakash K. Aithal et al. · 0 citations
Book Open access Aug 2026

Reinforcement Learning with Scenario-Context Rollout in Portfolio Management

This work proposes Scenario-Context Rollout (SCR), a macroeconomics-guided feedback mechanism to produce a distribution of next-day joint returns under potential economic shocks, and theoretically analyze this problem and shows that combining scenario-scored rewards with tape-realized transitions induces a hybrid fixed point.

Vanya Priscillia Bendatu, Yao Lu · 0 citations
Open access Jul 2026

AUTOMATING PORTFOLIO MANAGEMENT USING MULTI-AGENT SYSTEM WITH DYNAMIC PROMPT OPTIMISATION AND FEEDBACK LOOPS

A closed-loop multi-agent decision framework that introduces prompt-level learning as a scalable alternative to full model retraining and highlights the potential of prompt-level adaptation for building robust and autonomous financial decision systems.

Kandarp Mukeshkumar Sharda, Aliyu Sani Sambo · 0 citations
Open access Jul 2026

Adaptive Portfolio Optimization Using MVF with Machine Learning Forecasting and Regime Switching: Evidence from LQ45 Stocks

Findings indicate that combining machine-learning-based predictive modelling with adaptive, regime-driven allocation enhances portfolio stability, mitigates extreme losses, and improves risk-return efficiency under dynamic emerging-market conditions.

Fadly Ramdhani, D. Saepudin · 0 citations
Review Open access Aug 2026

A survey on LLM-enhanced reinforcement learning in financial markets

The integration of Large Language Models (LLMs) with Reinforcement Learning (RL) for financial decision-making has grown rapidly in recent years, yet the literature remains fragmented and lacks systematic comparison across methods. In this survey we analyze 34 core studies (2023–2026), selected through a multi-stage process involving 84 initial candidates and 46 full-text reviews, and propose a three-paradigm taxonomy (feature-based, auxiliary-based, and policy-based) based on the functional role of LLMs within the RL pipeline. Analysis of these integration paradigms reveals an emergent architectural trade-off: while tighter policy-based coupling theoretically offers deeper contextual reasoning, it frequently introduces significant computational overhead and training instability. Conversely, simpler feature-based integration provides superior scalability and stability, though often at the expense of representational depth. Given the current benchmark fragmentation, the reported performance gains across these studies remain difficult to validate universally across different asset classes. Critical gaps identified include the insufficient handling of data leakage and look-ahead bias, standardized benchmarks, and limited alignment with regulatory frameworks such as MiFID II and the EU AI Act.

Ghusoon Hadi al-Aldaffaie, Alireza Taheri, Amirfarhad Farhadi et al. · 0 citations
Conference Jul 2026

Portfolio Optimization Under Varying Market Regimes: A Comparative Study Using ANN Forecasts on MAI Stocks

This paper explores the effectiveness of integrating Artificial Neural Network (ANN) based return forecasts into three portfolio optimization frameworks: Equal-Weight (EW), Mean-Variance (MV), and Black-Litterman (BL), under varying market regimes, using 67 stocks listed on Thailand’s Market for Alternative Investment (MAI) as a case study. Portfolios are constructed using ANN predictions and tested across pre-COVID stable conditions (2019) and a volatile period (2020–2024), with an additional evaluation of rebalanced portfolios on 2024 data. Results indicate that BL achieves the highest risk-adjusted returns under stable conditions, while EW outperforms significantly during high volatility, driven largely by outlier stock performance during the COVID-19 recovery. Rebalancing with updated ANN forecasts did not guarantee improved performance, highlighting both the promise and limitations of machine learning-enhanced optimization in emerging markets.

Pema Yangchen, Rujira Chaysiri · 0 citations