Skip to content
Book Open access

TransX: Scaling Transformer-based Recommendation via Behavioral and Serving Stream Crossings

Jul 2026 · Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 · pp. 8312-8323 · 0 citations · 39 references
Computer Science

TL;DR

TransX is a production-oriented encoder–decoder architecture that reformulates recommendation as a sequence-to-sequence action transduction problem that decouples behavior-stream modeling from serving-event modeling and conditions next-action decoding on scalable cross-attention between nearline behavior encodings and real-time serving representations.

Abstract

Modern industrial recommender systems (RecSys) increasingly adopt Transformer-based sequence models, with an emerging paradigm that frames recommendation as next-token prediction over a unified monolithic user sequence. However, collapsing heterogeneous data sources -- such as long-term user behaviors and real-time serving events -- into a single monolithic token stream that obscures their distinct causal roles and temporal characteristics, leading to inefficient modeling and elevated training and serving costs. We propose TransX, a production-oriented encoder–decoder architecture that reformulates recommendation as a sequence-to-sequence action transduction problem. TransX explicitly decouples behavior-stream modeling from serving-event modeling and conditions next-action decoding on scalable cross-attention between nearline behavior encodings and real-time serving representations. To enable low-latency, high-QPS deployment, TransX is co-designed with an amortized serving strategy that combines incremental behavior encoding with per-request key–value caching, rendering serving latency insensitive to behavior sequence length. Extensive offline experiments and large-scale online A/B tests on LinkedIn's recommender systems show that TransX consistently outperforms state-of-the-art DLRMs and sequential baselines, and delivers substantial CTR lift (+6.0%) and conversion gain (+4.4%) while maintaining serving costs comparable to existing production models where our co-designed serving strategy reduces online computation by approximately 80%.

Read PDF

Similar papers

Book Open access Sep 2026

Transformer-based Sequential Recommender Systems

Transformer-based models have become the cornerstone of sequential recommendation, yet they are often perceived either as rigid engineering recipes or as a collection of disconnected architectures. This tutorial demystifies these systems by centering on a provocative guiding question: “What is not sequential recommenda...

J. Lichtenberg, A. V. Petrov · 0 citations
#artificial intelligence Preprint Sep 2026

From Language to Behavior: Scaling Sequence Transformers for Industrial Recommendation Ranking with Rec-Native Designs

ReST is proposed, a recommendation-native Transformer scaling framework that achieves higher accuracy and scales more consistently along sequence length, depth, and width, where LLM-style Transformer blocks saturate.

Jie Chen, Xiang-Qia Yu, Yan-Chao Lian et al. · 0 citations
Book Open access Sep 2026

DP-Rec: Towards Dynamic Patching for Efficient Long-Sequence Recommendation

Inspired by the recent success of Byte Latent Transformer, DP-Rec is proposed, a dynamic latent patching architecture for recommendation that scales effectively to long sequences and achieves a superior efficiency–accuracy trade-off over both non-compressed and fixed-size compression baselines.

Dwipam Katariya, T. Caputo, Akshat Shreemali et al. · 0 citations
Book Open access Sep 2026

UniTraj: Cross-Domain Long-Sequence Modeling for Commercial Recommendation

UniTraj, a practical framework that extends sequence construction beyond the advertising domain by incorporating behaviors from content-consumption scenarios, forming unified commercial trajectories across domains and scenarios, is proposed and deployed in a large-scale online advertising system.

Xian Hu, Ming Yue, Zhi-Xiang Feng et al. · 0 citations
Preprint Aug 2026

OneModel: A Unified Foundation for Platform-Scale Multi-Scenario Ranking

Platform-scale recommender systems often span multiple business streams such as organic recommendation, advertising, and merchant services, where user behaviors form a continuous cross-stream trajectory. Maintaining separate ranking systems fragments user representations and increases engineering cost. We propose \text...

Yinqi Zhang, Pei-Yu Hu, Yuntian Tang et al. · 1 citation
Book Open access Sep 2026

TransAct V2: Production System for Lifelong User Sequence Modeling at Scale

This work presents TransAct V2, a production ranking model deployed at Pinterest serving over 630 million users, that addresses challenges through a dual-path training-serving architecture with request-level deduplication and int8 quantization and provides a reproducible blueprint for practitioners building large-scale...

Xue Xia, S. Joshi, Kousik Rajesh et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.