Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

Dynamic portfolio optimization under uncertainty: A penalty-function-based neural solver

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

Dynamic portfolio optimization under uncertainty: A penalty-function-based neural solver

Similar Papers
  • Conference Article
  • Cite Count Icon 2
  • 10.1109/ccdc.2015.7162075
Dynamic mean-variance portfolio optimization with noshorting constraint and correlated returns
  • May 1, 2015
  • Chong Wei + 1 more

Rich empirical evidence shows the dependent relationship of the returns of risky assets between different time periods. However, current literature on the dynamic portfolio optimization model mainly adopts the assumption of independence of the returns. Due to the regulation of the financial market, the restriction of the short-selling is another indispensable factor in the portfolio management model. In this work, we study the discrete time dynamic mean-variance portfolio optimization model with noshorting constraint and the correlated returns of the risky assets. We adopt a formulation with a general structure of correlation, which enables us to better matching our model with real markets. By using the stochastic dynamic programming, we solve the problem analytically and derive the explicit portfolio policy, which is a piecewise affine function of the current wealth.

  • Research Article
  • Cite Count Icon 27
  • 10.1080/14697688.2018.1524155
Dynamic portfolio optimization with liquidity cost and market impact: a simulation-and-regression approach
  • Nov 14, 2018
  • Quantitative Finance
  • Rongju Zhang + 5 more

We present a simulation-and-regression method for solving dynamic portfolio optimization problems in the presence of general transaction costs, liquidity costs and market impact. This method extends the classical least squares Monte Carlo algorithm to incorporate switching costs, corresponding to transaction costs and transient liquidity costs, as well as multiple endogenous state variables, namely the portfolio value and the asset prices subject to permanent market impact. To handle endogenous state variables, we adapt a control randomization approach to portfolio optimization problems and further improve the numerical accuracy of this technique for the case of discrete controls. We validate our modified numerical method by solving a realistic cash-and-stock portfolio with a power-law liquidity model. We identify the certainty equivalent losses associated with ignoring liquidity effects, and illustrate how our dynamic optimization method protects the investor's capital under illiquid market conditions. Lastly, we analyze, under different liquidity conditions, the sensitivities of certainty equivalent returns and optimal allocations with respect to trading volume, stock price volatility, initial investment amount, risk aversion level and investment horizon.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 2
  • 10.1007/s10479-025-06649-x
A reinforcement learning approach to dynamic portfolio optimization
  • Jun 25, 2025
  • Annals of Operations Research
  • Maximilian Gollart + 1 more

This paper intends to bridge the gap between traditional and machine learning (ML) methods for dynamic portfolio optimization. We consider an investor who maximizes his utility from terminal wealth by dynamically allocating between risky and risk-free assets over time. Departing from the Deep Deterministic Policy Gradient (DDPG) algorithm by Lillicrap et al. (2016), we build a model-free reinforcement learning (RL) approach that is capable of deriving approximately optimal investment policies without any knowledge about the underlying market dynamics. It is agnostic to rebalancing frequency, allows for easy implementation of allocation constraints and requires low computational effort. For testing our algorithm we benchmark it against the theoretically optimal solution. Considering a realistic market model that is still theoretically feasible, enables us to show in detail the inherent connections between RL and dynamic portfolio optimization. Based on that we discuss the economic interpretability of this ML approach and the real-world applicability of such ML methods in general.

  • Supplementary Content
  • 10.25560/68488
Dynamic portfolio optimization with credit risk
  • Jan 1, 2018
  • Spiral (Imperial College London)
  • Longjie Jia

Credit risk, which considers the risk of loss resulting from a counterparty's failure to meet contractual obligations, was not properly studied in the literature of dynamic portfolio optimization problem until the 2008-2009 financial crisis. This thesis is devoted to the dynamic portfolio optimization with credit risk. Two main topics are studied in this thesis. In the first topic, we consider a utility maximization problem with defaultable stocks and looping contagion risk. We assume that the default intensity of one company depends on the stock prices of itself and other companies, and the default of the company induces immediate drops in the stock prices of the surviving companies. Under such looping contagion risk framework, we prove that the value function is the unique viscosity solution of the HJB equation. We also perform some numerical tests to compare and analyse the statistical distributions of the terminal wealth of log utility and power utility based on two strategies, one using the full information of intensity process and the other a proxy constant intensity process. The numerical tests confirm that modeling looping contagion risk properly is important, especially in financial distressed period. The second topic is on dynamic portfolio optimization with contingent convertible (CoCo) bond. As a new type of hybrid product, CoCo bond has the interesting feature that converting from debt to equity is contingent. We model the conversion of CoCo bond by reduced-form approach and assume that the conversion intensity is a deterministic function of the coupon rate and the issuing bank's stock price. Theoretically, we construct the viscosity solution representation between the value function and the corresponding HJB equation. Practically, we compare the performance between investing into CoCo bond and the issuing bank's stock. We analyse the statistical distributions of terminal wealth of log utility and power utility based on these two investment choices. Our simulation results show that, the CoCo bond holders bear more loss than equity holders if conversion occurs. However, investing into CoCo bond gets more profit (mean) while bearing less market risk (volatility) as long as conversion does not occur.

  • Conference Article
  • Cite Count Icon 1
  • 10.1109/ccdc.2016.7531149
Dynamic mean-LPM portfolio optimization under the mean-reverting market
  • May 1, 2016
  • Yiwei Niu + 1 more

This paper studies the dynamic mean-LPM portfolio optimization problem with mean-reverting market. Rich empirical evidence shows that the stock return exhibits certain degree of mean-reverting property. However, the study of dynamic mean-risk model with mean-reverting market is rare in the literature. Due to the complicated constraints in this kind of dynamic optimization model, it is hard to use the stochastic control approach directly. Instead, under our market setting, we are able solve this problem semi-analyticall by using the martingale approach. Once the optimal terminal wealth is identified, we use some numerical approach and the Monte Carlo method to find the optimal wealth process and portfolio process for given state. Our method provide investor a simple tool to deal with the sophisticated model for guiding the portfolio management.

  • Research Article
  • Cite Count Icon 15
  • 10.1007/s001860000049
Price systems constructed by optimal dynamic portfolios
  • Aug 17, 2000
  • Mathematical Methods of Operations Research (ZOR)
  • Manfred Schäl

The paper studies connections between arbitrage and utility maximization in a discrete-time financial market. The market is incomplete. Thus one has several choices of equivalent martingale measures to price contingent claims. Davis determines a unique price for a contingent claim which is based on an optimal dynamic portfolio by use of a `marginal rate of substitution' argument. Here conditions will be given such that this price is determined by a martingale measure and thus by a consistent price system. The underlying utility function U is defined on the positive half-line. Then dynamic portfolios are admissible if the terminal wealth is positive. In case of the logarithmic utility function, the optimal dynamic portfolio is the numeraire portfolio.

  • Research Article
  • Cite Count Icon 10
  • 10.1016/j.iref.2023.04.022
Central bank asset purchases, banks’ risky security holdings and profitability: Macro and micro evidence from Japan and the U.S.
  • Apr 26, 2023
  • International Review of Economics & Finance
  • Ling Wang

Central bank asset purchases, banks’ risky security holdings and profitability: Macro and micro evidence from Japan and the U.S.

  • Research Article
  • Cite Count Icon 11
  • 10.1007/s40747-025-01884-y
A novel multi-agent dynamic portfolio optimization learning system based on hierarchical deep reinforcement learning
  • May 29, 2025
  • Complex & Intelligent Systems
  • Ruoyu Sun + 4 more

Deep reinforcement learning (DRL) has been extensively used to address portfolio optimization problems. DRL agents acquire knowledge and make decisions through unsupervised interactions with their environment without requiring explicit knowledge of the joint dynamics of portfolio assets. Among these DRL algorithms, the combination of actor-critic algorithms and deep function approximators is the most widely used DRL algorithm. Here, we find that training the DRL agent using the actor-critic algorithm and deep function approximators may lead to scenarios where the improvement in the DRL agent's risk-adjusted profitability is insignificant. We argue that such situations primarily arise from the following two problems: sparsity in positive reward and the curse of dimensionality. These limitations prevent DRL agents from comprehensively learning asset price change patterns in the training environment. As a result, the DRL agents cannot effectively explore the dynamic portfolio optimization policy to improve the risk-adjusted profitability in the training process. To address these problems, we propose a novel multi-agent learning system based on the hierarchical deep reinforcement learning (HDRL) algorithmic framework in this research. Under this framework, the agents work together as a learning system for portfolio optimization. Specifically, by designing an auxiliary agent that works together with the executive agent for optimal policy exploration, the learning system can focus on exploring the policy with higher risk-adjusted return in the action space with positive return and low variance. The performance of the proposed learning system is evaluated using a portfolio of 29 stocks from the Dow Jones index in four different experiments. In the training process, the objective functions of the actor and critic both ultimately achieve stable convergence in the training process. The risk-adjusted profitability of our learning system in the training environment is significantly improved. Hence, we prove that the policies executed by our learning system in out-sample experiments originate from the DRL agents' comprehensive learning of asset price change patterns in the training environment. Furthermore, we find that adopting the auxiliary agent and HDRL training algorithm can efficiently overcome the issue of the curse of dimensionality and improve the training efficiency in the positive reward sparse environment. In each back-test experiment, the proposed learning system is compared to sixteen traditional strategies and ten strategies based on machine learning algorithms in the performance of profitability and risk control ability. The empirical results in the four evaluation experiments demonstrate the efficacy of our learning system, which outperforms all other strategies by at least 8.2% in terms of Sharpe ratio, Sorino ratio, and Calmar ratio. This indicates that the policies learned in the training environment can exhibit excellent generalization ability in the back-testing experiments.

  • Research Article
  • Cite Count Icon 34
  • 10.1016/j.eswa.2022.118739
Dynamic portfolio optimization with inverse covariance clustering
  • Sep 13, 2022
  • Expert Systems with Applications
  • Yuanrong Wang + 1 more

Market conditions change continuously. However, in portfolio investment strategies, it is hard to account for this intrinsic non-stationarity. In this paper, we propose to address this issue by using the Inverse Covariance Clustering (ICC) method to identify inherent market states and then integrate such states into a dynamic portfolio optimization process. Extensive experiments across three different markets, NASDAQ, FTSE and HS300, over a period of ten years, demonstrate the advantages of our proposed algorithm, termed Inverse Covariance Clustering-Portfolio Optimization (ICC-PO). The core of the ICC-PO methodology concerns the identification and clustering of market states from the analytics of past data and the forecasting of the future market state. It is therefore agnostic to the specific portfolio optimization method of choice. By applying the same portfolio optimization technique on a ICC temporal cluster, instead of the whole train period, we show that one can generate portfolios with substantially higher Sharpe Ratios, which are statistically more robust and resilient with great reductions in the maximum loss in extreme situations. This is shown to be consistent across markets, periods, optimization methods and selection of portfolio assets.

  • Research Article
  • Cite Count Icon 12
  • 10.1002/int.22870
Dynamic portfolio optimization using technical analysis‐based clustering
  • Mar 27, 2022
  • International Journal of Intelligent Systems
  • Ahmad Zaman Khan + 1 more

An accurate prediction of asset prices is perhaps the biggest challenge of any study in portfolio optimization. Asset prices are affected by several random and nonrandom factors, which makes them difficult to forecast. This paper proposes a two-phase dynamic portfolio optimization approach. In the first phase, assets are clustered into buy, sell, and hold groups using technical indicators. We provide a methodology to integrate the investor attitude (optimistic, pessimistic, or neutral) during the clustering phase. In the second phase, we input the clustered groups into a portfolio optimization model to obtain the optimum asset allocations. We use coherent fuzzy numbers to model the asset returns to integrate the investor attitude in this phase. The optimization model is solved using a genetic algorithm. The portfolios are rebalanced at regular intervals as new data becomes available. We illustrate the proposed methodology on a 100-asset problem of the US stock market. We analyze the real-world performance of the obtained portfolios. We compare the performance of the proposed approach with the mean–variance model, and other portfolios, such as the naïve portfolio and the NASDAQ-100 index.

  • Research Article
  • Cite Count Icon 24
  • 10.1016/j.ejor.2014.12.040
Dynamic portfolio optimization with transaction costs and state-dependent drift
  • Dec 31, 2014
  • European Journal of Operational Research
  • Jan Palczewski + 3 more

Dynamic portfolio optimization with transaction costs and state-dependent drift

  • Research Article
  • Cite Count Icon 33
  • 10.1007/s10479-018-2991-z
Time-consistent risk-constrained dynamic portfolio optimization with transactional costs and time-dependent returns
  • Aug 1, 2018
  • Annals of Operations Research
  • Davi Valladão + 2 more

Dynamic portfolio optimization has a vast literature exploring different simplifications by virtue of computational tractability of the problem. Previous works provide solution methods considering unrealistic assumptions, such as no transactional costs, small number of assets, specific choices of utility functions and oversimplified price dynamics. Other more realistic strategies use heuristic solution approaches to obtain suitable investment policies. In this work, we propose a time-consistent risk-constrained dynamic portfolio optimization model with transactional costs and Markovian time-dependence. The proposed model is efficiently solved using a Markov chained stochastic dual dynamic programming algorithm. We impose one-period conditional value-at-risk constraints, arguing that it is reasonable to assume that an investor knows how much he is willing to lose in a given period. In contrast to dynamic risk measures as the objective function, our time-consistent model has relatively complete recourse and a straightforward lower bound, considering a maximization problem. We use the proposed model for approximately solving: (i) an illustrative problem with 3 assets and 1 factor with an autoregressive dynamic; (ii) a high-dimensional problem with 100 assets and 5 factors following a discrete Markov chain. In both cases, we empirically show that our approximate solution is near-optimal for the original problem and significantly outperforms selected (heuristic) benchmarks. To the best of our knowledge, this is the first systematic approach for solving realistic time-consistent risk-constrained dynamic asset allocation problems.

  • Research Article
  • Cite Count Icon 2
  • 10.52396/justc-2022-0072
A new deep reinforcement learning model for dynamic portfolio optimization
  • Jan 1, 2022
  • JUSTC
  • Weiwei Zhuang + 2 more

There are many challenging problems for dynamic portfolio optimization using deep reinforcement learning, such as the high dimensions of the environmental and action spaces, as well as the extraction of useful information from a high-dimensional state space and noisy financial time-series data. To solve these problems, we propose a new model structure called the complete ensemble empirical mode decomposition with adaptive noise (CEEMDAN) method with multi-head attention reinforcement learning. This new model integrates data processing methods, a deep learning model, and a reinforcement learning model to improve the perception and decision-making abilities of investors. Empirical analysis shows that our proposed model structure has some advantages in dynamic portfolio optimization. Moreover, we find another robust investment strategy in the process of experimental comparison, where each stock in the portfolio is given the same capital and the structure is applied separately.

  • Research Article
  • 10.33395/sinkron.v9i3.14868
Hybrid Genetic Algorithm for Dynamic Portfolio Optimization Problems
  • Aug 10, 2025
  • sinkron
  • Sarah Ayatun Nufus + 2 more

Dynamic portfolio optimization is a complex problem due to continuous changes in market conditions, demanding algorithms capable of effective adaptation. Genetic Algorithms (GA) are often used for optimization problems but may face limitations in convergence speed and solution precision. This research aims to develop and evaluate a Hybrid Genetic Algorithm (HGA) that integrates GA with the Hill Climbing local search method, and to compare its performance against standard GA in solving dynamic portfolio optimization problems with the objective of maximizing the Sharpe Ratio. A series of simulation-based experiments were conducted by varying key algorithmic and dynamic environment parameters. Simulation results indicate that HGA generally has significant potential to improve performance compared to standard GA. Consistently, HGA successfully achieved superior solution quality, both in terms of Offline Performance Solution Quality and Overall Best Fitness. Regarding robustness to dynamic changes, HGA also demonstrated a smaller impact from performance degradation and a more promising recovery capability after market environment changes. Although HGA's superiority in convergence speed is not always absolute and the implementation of Hill Climbing adds to the computational time per generation, the improvement in solution quality and robustness offered in many configurations can be considered a worthwhile trade-off, especially for complex dynamic portfolio optimization problems. These findings support the hypo that hybridizing GA with local search can provide a positive contribution, noting that careful parameter tuning is crucial for maximizing HGA's potential.

  • Book Chapter
  • Cite Count Icon 6
  • 10.1007/978-3-030-41068-1_10
Applications of Reinforcement Learning
  • Jan 1, 2020
  • Matthew F Dixon + 2 more

This chapter considers real-world applications of reinforcement learning in finance, as well as further advances in the theory presented in the previous chapter. We start with one of the most common problems of quantitative finance, which is the problem of optimal portfolio trading in discrete time. Many practical problems of trading or risk management amount to different forms of dynamic portfolio optimization, with different optimization criteria, portfolio composition, and constraints. This chapter introduces a reinforcement learning approach to option pricing that generalizes the classical Black–Scholes model to a data-driven approach using Q-learning. It then presents a probabilistic extension of Q-learning called G-learning and shows how it can be used for dynamic portfolio optimization. For certain specifications of reward functions, G-learning is semi-analytically tractable and amounts to a probabilistic version of linear quadratic regulators (LQR). Detailed analyses of such cases are presented, and show their solutions with examples from problems of dynamic portfolio optimization and wealth management.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant