Jul 2026· IEEE Transactions on Neural Networks and Learning Systems· Vol PP, pp. 1-14· 0 citations
Medicine
TL;DR
Ablation studies comparing the AHRL-PM model to a synchronous model and single-layer reinforcement learning (RL) approaches validated its superior performance in terms of profitability and risk-adjusted returns, while also highlighting significant reductions in training time and appropriate portfolio weight adjustments to respond to market dynamics and uncertainties, underscoring the model's efficiency and practical applicability.
Abstract
Effective portfolio management (PM) is a cornerstone of financial strategy, yet it is often challenged by the uncertainties and the high dimensionality of market data. Traditional PM techniques, whether model-based or model-free, frequently fail to address these complexities, resulting in suboptimal asset allocation. This study introduces an asynchronous hierarchical reinforcement learning (AHRL-PM) framework aimed at enhancing active PM. The proposed AHRL-PM approach employs a two-tiered agent system to navigate market complexities. The first-layer agent selects top-performing stocks monthly based on alpha factors, while the second-layer agent adjusts asset weights using market fundamentals and technical indicators, dynamically rebalancing the portfolio in response to market shifts on a daily basis. We rigorously evaluated our AHRL-PM framework through extensive experiments in six globally representative markets. Performance was assessed using two key metrics: annualized return (AR) and Sharpe ratio (SR). Our model consistently outperformed traditional benchmarks on both metrics across all the markets, demonstrating robustness and versatility in PM. Furthermore, ablation studies comparing the AHRL-PM model to a synchronous model and single-layer reinforcement learning (RL) approaches validated its superior performance in terms of profitability and risk-adjusted returns, while also highlighting significant reductions in training time and appropriate portfolio weight adjustments to respond to market dynamics and uncertainties, underscoring the model's efficiency and practical applicability.
The systematic development of single-agent to multi-agent ensemble systems shows great improvements in algorithmic and architecture of DRL-based portfolio management, and the research in the future focuses on the importance of explainable AI integration, meta-learning market regime adaptation, and consistent evaluation systems in reproducible research.
Aditi Kumar Rout, U. D. Acharya, Prakash K. Aithal et al.· Discover Computing· 0 citations
This work proposes Scenario-Context Rollout (SCR), a macroeconomics-guided feedback mechanism to produce a distribution of next-day joint returns under potential economic shocks, and theoretically analyze this problem and shows that combining scenario-scored rewards with tape-realized transitions induces a hybrid fixed point.
Vanya Priscillia Bendatu, Yao Lu· Proceedings of the 32nd ACM...· 0 citations
A closed-loop multi-agent decision framework that introduces prompt-level learning as a scalable alternative to full model retraining and highlights the potential of prompt-level adaptation for building robust and autonomous financial decision systems.
Kandarp Mukeshkumar Sharda, Aliyu Sani Sambo· NLP & Big Data· 0 citations
The integration of Large Language Models (LLMs) with Reinforcement Learning (RL) for financial decision-making has grown rapidly in recent years, yet the literature remains fragmented and lacks systematic comparison across methods. In this survey we analyze 34 core studies (2023–2026), selected through a multi-stage process involving 84 initial candidates and 46 full-text reviews, and propose a three-paradigm taxonomy (feature-based, auxiliary-based, and policy-based) based on the functional role of LLMs within the RL pipeline. Analysis of these integration paradigms reveals an emergent architectural trade-off: while tighter policy-based coupling theoretically offers deeper contextual reasoning, it frequently introduces significant computational overhead and training instability. Conversely, simpler feature-based integration provides superior scalability and stability, though often at the expense of representational depth. Given the current benchmark fragmentation, the reported performance gains across these studies remain difficult to validate universally across different asset classes. Critical gaps identified include the insufficient handling of data leakage and look-ahead bias, standardized benchmarks, and limited alignment with regulatory frameworks such as MiFID II and the EU AI Act.
Ghusoon Hadi al-Aldaffaie, Alireza Taheri, Amirfarhad Farhadi et al.· Discover Artificial Intellig...· 0 citations
This paper explores the effectiveness of integrating Artificial Neural Network (ANN) based return forecasts into three portfolio optimization frameworks: Equal-Weight (EW), Mean-Variance (MV), and Black-Litterman (BL), under varying market regimes, using 67 stocks listed on Thailand’s Market for Alternative Investment (MAI) as a case study. Portfolios are constructed using ANN predictions and tested across pre-COVID stable conditions (2019) and a volatile period (2020–2024), with an additional evaluation of rebalanced portfolios on 2024 data. Results indicate that BL achieves the highest risk-adjusted returns under stable conditions, while EW outperforms significantly during high volatility, driven largely by outlier stock performance during the COVID-19 recovery. Rebalancing with updated ANN forecasts did not guarantee improved performance, highlighting both the promise and limitations of machine learning-enhanced optimization in emerging markets.
Pema Yangchen, Rujira Chaysiri· International Conference on...· 0 citations