Skip to content

A comparison between Thompson sampling and greedy algorithm in portfolio selection

TL;DR

The findings suggest that in standard passive investment settings, the additional complexity of exploration-based reinforcement learning is not justified, and simpler estimation-based approaches deliver comparable performance.

Abstract

Portfolio optimization is a decision-making problem that allocates assets to achieve an optimal return–risk trade-off. The classical framework assumes that return and risk parameters are known; in practice, however, these parameters must be estimated from data, introducing estimation risk beyond the original model. Integrating parameter learning with portfolio choice naturally leads to a sequential decision-making framework based on reinforcement learning. This study examines sequential portfolio selection under parameter uncertainty within a Bayesian portfolio framework. It evaluates whether exploration through Thompson sampling improves performance relative to an exploration-free greedy strategy when asset returns are observable and portfolio weights do not influence the return-generating process. The analysis includes a simulation study and an empirical application using daily returns of eight stocks from the SET100 index. Portfolio weights are updated sequentially as posterior beliefs about expected returns evolve over time. Performance is assessed using cumulative return, Sharpe ratio, and regret. Simulation results show that Thompson sampling does not outperform the greedy algorithm across performance metrics; both approaches generate nearly identical outcomes. Empirical results using real data indicate that Thompson sampling occasionally slightly outperforms the greedy algorithm, but the differences are not economically significant. Overall, the findings suggest that in standard passive investment settings, the additional complexity of exploration-based reinforcement learning is not justified. Simpler estimation-based approaches deliver comparable performance.

View source

Similar papers

Open access Aug 2026

MEAN-VARIANCE PORTFOLIO OPTIMIZATION FOR EMDE AND MTDL STOCKS: A MARKOWITZ APPROACH

Constructing an optimal portfolio is a crucial step for investors in balancing the trade-off between expected return and investment risk. This study aims to construct an optimal portfolio comprising two stocks, EMDE and MTDL, by applying the Markowitz mean-variance model to minimize return variance at a specific return...

Zahra Rohadatul Aisylah, Ferdiansyah Saputra, Arief Surya Lesmana et al. · 0 citations
Open access Sep 2026

Portfolio Optimization Using Modern Portfolio Theory in Investment Management

Portfolio optimization is a fundamental aspect of investment management that focuses on constructing a portfolio capable of delivering the highest possible return while minimizing investment risk. Modern Portfolio Theory (MPT), introduced by Harry Markowitz, provides a quantitative framework for selecting an optimal co...

K. Naveen, Amita Johar, T. Meghana · 0 citations
Preprint Sep 2026

Active Portfolio Management in Concentrated Equity Markets

The equal-weighted portfolio is a passive, rule-based strategy that has historically been difficult to outperform, delivering higher returns than the capitalization-weighted"market"benchmark across many markets and periods. Stochastic portfolio theory (SPT) reveals that this relative performance is regime dependent, wi...

Brian Ceco, Xiao-Fei Shi, Ting-Kam Leonard Wong · 0 citations
Conference Aug 2026

Automation of Black-Litterman Asset Allocation Model Using Machine Learning for Algorithmic Stock Trading

Markowitz's mean-variance portfolio optimization model, while foundational in modern finance, is known for its sensitivity to input parameters and tendency toward estimation error maximization, often resulting in overly concentrated portfolios. The Black-Litterman model presents a refined alternative by leveraging Baye...

Vishvajeet Upadhyay, Porinita Banerjee, Shivaditya Upadhyay · 0 citations
Open access Aug 2026

Optimal Portfolio Strategy for A Sensitive Investor in A Dynamic Financial Market

This paper investigates optimal portfolio choice for a risk-averse investor who is operating in a financial market characterized by continuous time usage and with explicit attention being paid to the investor's sensitivity to market movements. The investor's preferences are described by a power utility function of c...

C. Achudume · 0 citations
Conference Aug 2026

MCDM-Guided Deep Reinforcement Learning for Dynamic Portfolio Optimization

When there is uncertainty, portfolio optimisation is the process of figuring out how to invest money in a group of assets. This is an important part of managing assets. Intellectual portfolio optimisation is now necessary in today's financial markets because the market is becoming more complicated and unstable. Current...

Anupama Pandey, Ayush Kumar Agrawal, Abhinav Shukla et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.