#machine learning
May 2026
Latent Order Bandits
A posterior-sampling algorithm is proposed and shown that both are competitive with full-prior latent bandits when same-state instances share reward parameters, and preferable to them when reward scales differ between instances with the same latent state.
Emil Carlsson, Newton Mwai, Fredrik D. Johansson
· arXiv.org · 0 citations