Skip to content

Challenges and Solutions for Bandits in the Wild: Warm-Started Mixture Bandits for Cross-Cohort Slate Recommendation

Sep 2026 · 0 citations · 46 references
Computer Science

TL;DR

This work proposes CohortMix-TS, a warm-started mixture bandit that learns latent user groups from earlier cohorts and uses available metadata to construct group-informed priors for new users, and shows how warm-start transfer and inventory-aware recommendations can support personalization for short-lived, repeatedly cold-starting cohorts.

Abstract

Many recommender services repeatedly encounter cold-start cohorts, where new users arrive with little or no interaction history. This creates two challenges: learning user preferences quickly from limited feedback and sustaining useful recommendations when each user has a finite catalog that can become repetitive or depleted over time. We propose CohortMix-TS, a warm-started mixture bandit that learns latent user groups from earlier cohorts and uses available metadata to construct group-informed priors for new users. Starting from these fixed priors, the model personalizes independently as feedback from each user becomes available. Session slates combine Thompson sampling with diversity and inventory-depletion controls. We evaluate CohortMix-TS through simulation, semi-synthetic experiments, and a 25-day randomized in-the-wild deployment with 713 registered participants in a Campus Games quiz application. Our evaluations show that cross-cohort transfer improves early recommendation quality and user-level regret, while inventory-aware slate construction helps prevent premature exhaustion of preferred items. In the field deployment, treatment users also showed a larger early-to-late change in correctness than users receiving random recommendations. Together, these results show how warm-start transfer and inventory-aware recommendations can support personalization for short-lived, repeatedly cold-starting cohorts.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

COPE: Continual Personalization of LLMs under Sparse User Feedback via User Embeddings and Self-Evaluation

Experiments show that COPE consistently outperforms strong training-free and training-based baselines under sparse feedback, and remains complementary to Retrieval-Augmented Prompting, and further analyses confirm COPE's reliable self-evaluation, meaningful preference patterns, stable general capabilities, and robustne...

Rui-Ke Cao, Fu-Gen Yao, Liang Dong et al. · 0 citations
Open access Aug 2026

Multimodal Interest-Shifting Sequence Recommendation with Offline Policy Optimization and Drift-Aware Representation Learning

The reported results indicate competitive public-data sequential ranking under the stated protocol, together with a reproducible and inference-safe evaluation design, and the practical value of the method lies in reaching, and in several comparisons, significantly exceeding, the accuracy of heavyweight multimodal-LLM-s...

Chang-Cheng Shao, Cheng Zeng, Xiao-Gang Ye et al. · 0 citations
#machine learning Preprint Oct 2026

Collaborative Personalized Preference Alignment for LLMs under Data Deficiency

Real-world users often exhibit highly heterogeneous preferences over multiple objectives for LLM responses. A lightweight aligner can tailor these responses to individual preferences, but scarce user-specific feedback makes personalized training difficult. Learning shared initializations across users can support few-sh...

Li-Yan Yang, Yi-Ge Yuan, Zhiqin Yang · 0 citations
Book Open access Sep 2026

In-Batch Negatives Can Silently Cripple LLM-Encoded Sequential Recommenders

LLM-based sequential recommenders increasingly train user/item encoders with in-batch InfoNCE negatives inherited from contrastive learning—convenient, but it silently caps the negative pool at batch size rather than catalog size whenever encoder cost bounds the batch. We report an early finding: a 15-negative in-batch...

Younggue Bae · 0 citations
Preprint Sep 2026

Recommender System as Slow and Fast Thinkers

Experiments on five real-world datasets show that \textsc{DS-Frame} consistently improves representative sequential recommendation backbones, with larger gains on challenging groups and effective accuracy--efficiency trade-offs, highlighting the potential of adaptive inference for more efficient and robust recommendati...

Zi-Chen Yuan, Xiao-Xuan Dong, Linkun Dai et al. · 0 citations
Book Open access Sep 2026

Support Gap: Selecting Fixed-K Candidate Sets for Retained Personalized Headroom

Two-stage recommender systems often must choose among same-budget candidate sets before expensive reranking or online evaluation, yet Recall@K, diversity, and coverage do not measure how much fine-state-contingent choice survives retrieval. We define retained personalized headroom as the value of choosing separately fo...

Teresa Zhang · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.