Aug 2026· International Conference on Information Security and Cryptology· pp. 1659-1664· 0 citations· 16 references
Abstract
Portfolio optimization is a fundamental problem in finance which has normally been addressed by mean-variance frameworks and their extensions. However, these methods rely on assumptions such as normally distributed returns and covariance estimates which often fail to capture the dynamics of real markets. Advances in machine learning have provided new tools for modelling decision-making and adapting to changing environments. This study analyzes a range of machine learning approaches to portfolio optimization, from predictive modelling with classical optimization to end-to-end reinforcement learning frameworks. We review methods such as deep neural networks for return forecasting and actor-critic algorithms (DDPG, PPO, SAC) for dynamic asset allocation. Studies in the literature review demonstrate that RL-based methods can outperform static strategies on metrics such as the Sharpe ratio. Regardless, challenges remain in terms of overfitting, interpretability, and scalability to large asset universes. By combining findings across different approaches, this study highlights the trade-offs between predictive and optimization hybrids and fully model-free RL and outlines future works in multimodal learning and risk-constrained optimization.
This paper presents the RL-DynTrade framework by using a cutting-edge deep reinforcement learning method, Proximal Policy Optimization (PPO), with a Deep Q Network (DQN) agent to dynamically adapt to changing risk-reward dynamics. PPO enables real-time, fine-grained, risk-reward adaptation via an actor-critic design wi...
Xi-Jing Ou, Jie Huang· International Journal of e-c...· 0 citations
Mean-variance portfolio optimization (MVO) is a central framework in data-driven asset management. A widely adopted approach is a two-stage framework that first predicts expected returns and then solves the optimization problem based on these predictions, with the predictive models trained by minimizing prediction erro...
A three-paradigm taxonomy (feature-based, auxiliary-based, and policy-based) based on the functional role of LLMs within the RL pipeline is proposed, which provides superior scalability and stability, though often at the expense of representational depth.
Ghusoon Hadi al-Aldaffaie, Alireza Taheri, Amirfarhad Farhadi et al.· Discover Artificial Intellig...· 0 citations
The results show that deep learning models perform best in highly efficient markets where signals are weak but consistent, and in moderately and least efficient markets, traditional strategies often achieve similar or better returns.
H. Sahu, Avishek Bhandari· Discover Artificial Intellig...· 0 citations
This study examines whether integrating machine learning-based return forecasting and dynamic covariance estimation into a mean-variance portfolio framework produces measurable improvements over a conventional benchmark. Three strategies are constructed and evaluated over a five-year out-of-sample window from January 2...
Gun Assavasopee· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.