Deep reinforcement learning for portfolio allocation: a comparative study on five Vietnamese equities
W. Saijai, Kansuda Pankwaen
Abstract
This study applies deep reinforcement learning (DRL) to multi-asset portfolio optimization in the Vietnamese stock market, aiming to evaluate the performance and stability of different DRL algorithms under emerging market conditions. Seven algorithms – A2C, proximal policy optimization (PPO), deep deterministic policy gradient, twin delayed deep deterministic policy gradient (TD3), soft actor-critic (SAC), truncated quantile critics (TQC) and RecurrentPPO – are trained on daily data from January 2018 to December 2024 and evaluated out-of-sample from January 2023 through September 2025 using five liquid equities (SBT.VN, BID.VN, CTG.VN, HPG.VN and VCB.VN). The state representation includes technical indicators (RSI, MACD, SMA, EMA, Bollinger Bands and OBV), as well as risk features such as rolling covariance and a turbulence index. PPO achieves the highest annual return (0.1666) and demonstrates the most stable performance. TD3 delivers comparable cumulative growth with higher variability. RecurrentPPO attains the highest Sharpe ratio (1.0461), highlighting the importance of temporal modeling. SAC and TQC produce more conservative but stable outcomes. This study does not propose a new DRL algorithm; instead, it provides a controlled and reproducible benchmarking framework for comparing multiple DRL models under identical conditions in an emerging market setting.
Source: semanticscholar · PDF
Read the AI summary and key takeaways for traders on WOBR Quant Research.