A CVaR-Constrained Safe Reinforcement Learning Framework With Action Repair for Practical Portfolio Optimization
Himanshu Choudhary, Arishi Orra, Manoj Thakur, Xiao-Zhi Gao, P. Sahu
Abstract
Portfolio optimization refers to allocating capital across multiple assets to attain an optimal balance between potential returns and associated risks under dynamic market conditions. While deep reinforcement learning (DRL) has shown significant promise in learning adaptive allocation strategies through continuous interaction with the market, most existing approaches struggle to incorporate practical investment constraints and robust risk measures. To address these limitations, this article introduces a constraint-aware and risk-sensitive DRL framework for portfolio optimization that operates within a constrained Markov decision process (CMDP) setting. The proposed model embeds conditional value-at-risk (CVaR) as a hard safety constraint using an augmented Lagrangian approach on the agent to prevent high-loss outcomes. Additionally, a repair mechanism is introduced to adjust the agent’s actions, which ensures the resulting portfolio weights remain feasible and compliant with real-world investment constraints. Our approach is evaluated using real-world stock data from multiple global indices and consistently outperforms existing DRL-based and traditional portfolio optimization strategies, achieving superior returns while strictly adhering to defined risk constraints.
Source: semanticscholar
Read the AI summary, key takeaways and discussion on WOBR Quant Research.