Skip to content
Conference

Advances in Portfolio Optimization from Mean-Variance to Reinforcement Learning

Aug 2026 · International Conference on Information Security and Cryptology · pp. 1659-1664 · 0 citations · 16 references

Abstract

Portfolio optimization is a fundamental problem in finance which has normally been addressed by mean-variance frameworks and their extensions. However, these methods rely on assumptions such as normally distributed returns and covariance estimates which often fail to capture the dynamics of real markets. Advances in machine learning have provided new tools for modelling decision-making and adapting to changing environments. This study analyzes a range of machine learning approaches to portfolio optimization, from predictive modelling with classical optimization to end-to-end reinforcement learning frameworks. We review methods such as deep neural networks for return forecasting and actor-critic algorithms (DDPG, PPO, SAC) for dynamic asset allocation. Studies in the literature review demonstrate that RL-based methods can outperform static strategies on metrics such as the Sharpe ratio. Regardless, challenges remain in terms of overfitting, interpretability, and scalability to large asset universes. By combining findings across different approaches, this study highlights the trade-offs between predictive and optimization hybrids and fully model-free RL and outlines future works in multimodal learning and risk-constrained optimization.

View source

Similar papers

#reinforcement learning Open access Sep 2026

Reinforcement Learning-Driven Dynamic Trading Strategies for Financial Markets

This paper presents the RL-DynTrade framework by using a cutting-edge deep reinforcement learning method, Proximal Policy Optimization (PPO), with a Deep Q Network (DQN) agent to dynamically adapt to changing risk-reward dynamics. PPO enables real-time, fine-grained, risk-reward adaptation via an actor-critic design wi...

Xi-Jing Ou, Jie Huang · 0 citations
#machine learning Preprint Sep 2026

Decision-Focused Learning for Mean-Variance Portfolio Optimization via KKT-Based Reformulation

Mean-variance portfolio optimization (MVO) is a central framework in data-driven asset management. A widely adopted approach is a two-stage framework that first predicts expected returns and then solves the optimization problem based on these predictions, with the predictive models trained by minimizing prediction erro...

Kensei Nosaka, Shunnosuke Ikeda, Yuichi Takano · 0 citations
Review Open access Aug 2026

A survey on LLM-enhanced reinforcement learning in financial markets

A three-paradigm taxonomy (feature-based, auxiliary-based, and policy-based) based on the functional role of LLMs within the RL pipeline is proposed, which provides superior scalability and stability, though often at the expense of representational depth.

Ghusoon Hadi al-Aldaffaie, Alireza Taheri, Amirfarhad Farhadi et al. · 0 citations
Open access Aug 2026

Benchmarking deep reinforcement learning and classical models for portfolio optimization across market efficiency regimes

The results show that deep learning models perform best in highly efficient markets where signals are weak but consistent, and in moderately and least efficient markets, traditional strategies often achieve similar or better returns.

H. Sahu, Avishek Bhandari · 0 citations

To boost or not to boost? XGBoost and DCC-GARCH integration for mean-variance portfolio optimization

This study examines whether integrating machine learning-based return forecasting and dynamic covariance estimation into a mean-variance portfolio framework produces measurable improvements over a conventional benchmark. Three strategies are constructed and evaluated over a five-year out-of-sample window from January 2...

Gun Assavasopee · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.