A Mechanism-Level Comparison of ETC, UCB, and Thompson Sampling Under Finite-Horizon Constraints
Abstract
. Multi-armed bandit (MAB) algorithms are commonly used in sequential decision tasks such as online recommendation and advertising. In real-world systems, algorithms often have limited time and data to learn from, and poor decisions can be costly. From a finite-horizon perspective, this paper presents a mechanism-level comparison of Explore-Then-Commit (ETC), Upper Confidence Bound (UCB), and Thompson Sampling (TS). The findings suggest ETC is highly sensitive to the length of the exploration phase and does not allow parameter adjustment once it enters the commit stage. UCB tends to incur relatively high exploration costs in short horizons and is also sensitive to parameter settings. In contrast, Thompson Sampling usually exhibits smoother behavior in the early stages and more stable performance under finite horizons, with less reliance on parameter tuning, although some randomness across runs remains. Overall, algorithm selection in real-world systems should consider early-stage behavior and risk under finite-horizon constraints.