Results show that a single jointly trained model can replace a mature multi-stage recommendation cascade while improving recommendation quality on live traffic.
Abstract
We introduce Sona, a single-model generative recommender for Yandex Music. In an online A/B test, Sona replaced the entire production cascade, comprising more than 15 candidate generators followed by pre-ranking and ranking models that consume hundreds of features, including signals from large transformer models such as Argus and target-attention scorers, while significantly improving key engagement metrics. The architecture of Sona unifies candidate generation and ranking around a shared user representation. Its encoder transforms the user's chronological sequence of logged engagement events into hidden states consumed by both the autoregressive decoder and the Ranking Module. The next-token-prediction and distillation objectives jointly update the encoder, coupling generation and ranking through the same user state. Neither Sona nor its Teacher Ranker uses hand-engineered features; both operate on logged event fields and learned item representations. In the final Sona configuration, the larger teacher supplies ranking targets during training but is absent from serving, leaving the encoder, decoder, and Ranking Module as a single deployed model. We evaluate Sona in an online A/B experiment using live traffic from My Vibe on smart speakers, one of Yandex Music's largest recommendation surfaces. Relative to the production control, Sona produced statistically significant uplifts of 4.53% in Active Users, the primary metric, 6.30% in Total Listening Time, and 11.42% in Likes. These effects were incremental to improvements retained from preceding deployments. The Active Users uplift was 2.35 times the increment previously delivered by Argus, the strongest model deployed on this surface before Sona. These results show that a single jointly trained model can replace a mature multi-stage recommendation cascade while improving recommendation quality on live traffic.
Gryphon-v2, a unified generate-and-rank architecture for end-to-end recommendation, and results support the practical viability of a generative retriever with a Ranking Module distilled from the Teacher Ranker as an end-to-end alternative to a production cascade.
Anna Lipkina, Daria Tikhonovich, Viktor Yanush et al.· 0 citations
We present suryaaseran1, our conversational music recommender for the TalkPlayData Challenge (ACM RecSys 2026) that treats each dialogue turn’s user reactions as a preference-elicitation signal for both retrieval and ranking. The dataset’s goal-progress feedback is delayed by one turn and includes rejected tracks logge...
Suryaa Veerabathiran Seran· Proceedings of the Workshop...· 1 citation· ⚡1
This diagnostic compares the full pipeline with reciprocal-rank fusion, ablate per-retriever features, and decompose ranking error into retrieval misses, reranking exclusions, and ranks-2–20 ordering loss and complement the leaderboard result with a diagnostic that fits on Train and evaluates the official Devset.
Ryohei Wakatsuki· Proceedings of the Workshop...· 1 citation
CUEES (Correlation-gUided Encoder Selection), a lightweight heuristic that estimates complementarity through task- and category-level Pearson correlations between encoders'performance profiles, scoring a candidate set from single-encoder evaluations alone--without fusion training during selection.
Pei-Jun Liao, Hung-Shin Lee, Wen-Ze Ren et al.· 0 citations
The RecSys Challenge 2026 Music-CRS (TalkPlay) task formalizes this as two coupled sub-problems: given dialogue history and user context, retrieve a ranked list of the top-20 tracks from the full, unrestricted catalog, and generate a response that justifies the recommendation while sustaining conversational coherence.
Simran Sundrani, Mohan Bhambhani· Proceedings of the Workshop...· 1 citation· ⚡1
Results from a large-scale A/B test comparing GenRec against the current production ranker model are reported, where it is shown that a GenRec model trained with substantially fewer Phase-2 labeled training examples and input signals can achieve statistically significant gains in offline and online metrics.
Ying Li, Shradha Sehgal, A. Rao et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.