Jul 2026· Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2· pp. 8312-8323· 0 citations· 39 references
Computer Science
TL;DR
TransX is a production-oriented encoder–decoder architecture that reformulates recommendation as a sequence-to-sequence action transduction problem that decouples behavior-stream modeling from serving-event modeling and conditions next-action decoding on scalable cross-attention between nearline behavior encodings and real-time serving representations.
Abstract
Modern industrial recommender systems (RecSys) increasingly adopt Transformer-based sequence models, with an emerging paradigm that frames recommendation as next-token prediction over a unified monolithic user sequence. However, collapsing heterogeneous data sources -- such as long-term user behaviors and real-time serving events -- into a single monolithic token stream that obscures their distinct causal roles and temporal characteristics, leading to inefficient modeling and elevated training and serving costs. We propose TransX, a production-oriented encoder–decoder architecture that reformulates recommendation as a sequence-to-sequence action transduction problem. TransX explicitly decouples behavior-stream modeling from serving-event modeling and conditions next-action decoding on scalable cross-attention between nearline behavior encodings and real-time serving representations. To enable low-latency, high-QPS deployment, TransX is co-designed with an amortized serving strategy that combines incremental behavior encoding with per-request key–value caching, rendering serving latency insensitive to behavior sequence length. Extensive offline experiments and large-scale online A/B tests on LinkedIn's recommender systems show that TransX consistently outperforms state-of-the-art DLRMs and sequential baselines, and delivers substantial CTR lift (+6.0%) and conversion gain (+4.4%) while maintaining serving costs comparable to existing production models where our co-designed serving strategy reduces online computation by approximately 80%.
Transformer-based models have become the cornerstone of sequential recommendation, yet they are often perceived either as rigid engineering recipes or as a collection of disconnected architectures. This tutorial demystifies these systems by centering on a provocative guiding question: “What is not sequential recommenda...
J. Lichtenberg, A. V. Petrov· Proceedings of the 20th ACM...· 0 citations
ReST is proposed, a recommendation-native Transformer scaling framework that achieves higher accuracy and scales more consistently along sequence length, depth, and width, where LLM-style Transformer blocks saturate.
Jie Chen, Xiang-Qia Yu, Yan-Chao Lian et al.· 0 citations
Inspired by the recent success of Byte Latent Transformer, DP-Rec is proposed, a dynamic latent patching architecture for recommendation that scales effectively to long sequences and achieves a superior efficiency–accuracy trade-off over both non-compressed and fixed-size compression baselines.
Dwipam Katariya, T. Caputo, Akshat Shreemali et al.· Proceedings of the 20th ACM...· 0 citations
UniTraj, a practical framework that extends sequence construction beyond the advertising domain by incorporating behaviors from content-consumption scenarios, forming unified commercial trajectories across domains and scenarios, is proposed and deployed in a large-scale online advertising system.
Xian Hu, Ming Yue, Zhi-Xiang Feng et al.· Proceedings of the 20th ACM...· 0 citations
Platform-scale recommender systems often span multiple business streams such as organic recommendation, advertising, and merchant services, where user behaviors form a continuous cross-stream trajectory. Maintaining separate ranking systems fragments user representations and increases engineering cost. We propose \text...
Yinqi Zhang, Pei-Yu Hu, Yuntian Tang et al.· 1 citation
This work presents TransAct V2, a production ranking model deployed at Pinterest serving over 630 million users, that addresses challenges through a dual-path training-serving architecture with request-level deduplication and int8 quantization and provides a reproducible blueprint for practitioners building large-scale...
Xue Xia, S. Joshi, Kousik Rajesh et al.· Proceedings of the 20th ACM...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.