Scalable Pontryagin-Guided Adjoint-to-Control Recovery for Constrained Dynamic Portfolio Choice
Jaegi Jeon, Jeonggyu Huh, Hyeng Keun Koo, Byung Hwa Lim
Abstract
We develop a scalable adjoint-to-control framework for continuous-time portfolio choice under smooth pointwise constraints. A feasible direct-policy-optimization (DPO) policy supplies rollouts; after training, fixed-latent open-loop BPTT (OL-BPTT) yields first- and second-order pathwise sensitivities, whose conditional projections produce adapted adjoint inputs. A nested antithetic common-random-number regression estimates the shifted wealth-row martingale input, and deployment solves the local generalized-Hamiltonian problem by an exact QP for quadratic-affine blocks or by a log barrier otherwise. We prove an OL-BPTT--PMP correspondence retaining orthogonal projection residuals, a local barrier--KKT approximation, and a performance-to-adjoint bridge under local quadratic growth. In an n=100 constrained Merton benchmark, the learned first adjoint has 0.46% mean relative error; under the analytical policy, the first adjoint, wealth curvature, and Brownian coefficient have nRMSEs of 0.031%, 0.035%, and 0.326%. In a terminal-only predictable-return CRRA benchmark, wealth homogeneity gives $P^{X\bullet,*}=D^2_{X\bullet}V$ and $ζ^{X,*}=0$. At the primary $512\times16$ projection budget, the full estimated-shift decoder has policy RMSE below $8.5\times10^{-3}$, while the benchmark-specific zero-shift oracle is below $2\times10^{-4}$. Across constraint, factor, barrier, and switching-region audits, recovery sharply reduces local KKT residuals and remains practical for portfolios with up to 100 risky assets.
Read the AI summary, key takeaways and discussion on WOBR Quant Research.