Sep 2026· Proceedings of the 20th ACM Conference on Recommender Systems· 0 citations· 17 references
TL;DR
This work presents TransAct V2, a production ranking model deployed at Pinterest serving over 630 million users, that addresses challenges through a dual-path training-serving architecture with request-level deduplication and int8 quantization and provides a reproducible blueprint for practitioners building large-scale sequential recommenders.
Abstract
Modeling lifelong user action sequences in CTR prediction faces three production challenges: computational complexity with quadratic transformer costs as sequences extend from hundreds to tens of thousands of actions, infrastructure overhead from storage and network costs that scale with both sequence length and number of candidates per request, and weak training supervision as sequence encoders sit far from CTR prediction heads. We present TransAct V2, a production ranking model deployed at Pinterest serving over 630 million users, that addresses these challenges through: (1) a dual-path training-serving architecture with request-level deduplication and int8 quantization achieving 1% logging cost, (2) fused Triton kernels with pinned memory management delivering 250 × p99 latency improvement, and (3) a Next Action Loss providing direct supervision for sequence modeling within CTR frameworks. Deployed in production, TransAct V2 demonstrates significant improvements in both engagement and recommendation diversity. To support industry adoption, we open-source our serving optimizations with comprehensive ablation studies. Our work provides a reproducible blueprint for practitioners building large-scale sequential recommenders.
ReST is proposed, a recommendation-native Transformer scaling framework that achieves higher accuracy and scales more consistently along sequence length, depth, and width, where LLM-style Transformer blocks saturate.
Jie Chen, Xiang-Qia Yu, Yan-Chao Lian et al.· 0 citations
UniTraj, a practical framework that extends sequence construction beyond the advertising domain by incorporating behaviors from content-consumption scenarios, forming unified commercial trajectories across domains and scenarios, is proposed and deployed in a large-scale online advertising system.
Xian Hu, Ming Yue, Zhi-Xiang Feng et al.· Proceedings of the 20th ACM...· 0 citations
KuaFu is a unified behavior-compression layer whose minimal unit is one behavior item, with fidelity-oriented four-stage training and layered intermediate evaluation, which matches or exceeds uncompressed single-task production models on all five headline metrics, raises per-GPU throughput by 37%-350%, and saves 190 GP...
Jia-Hao Hui, Lin Zhu, Yi-Sheng Hu et al.· 0 citations
Ultra-long user history modeling has been a highly effective approach in modern industrial recommendation systems, with most works heavily utilizing search-based methods and summarization based methods to map massive user interaction logs into latent user representation with algorithm designs to address the scaling cha...
Yuan-Zheng Lin, Diego Uribe Mora, Yuan Shao et al.· Proceedings of the 20th ACM...· 0 citations
This work presents SAGA, a generative action embedding model that encodes multi-surface user interaction sequences across a Financial Service organization's ecosystems, from checkout, peer-to-peer (P2P) transactions, in-app engagement, email to account actions, into a unified user representation for downstream recommen...
Tsz-Fung Pang, Po‐Jui Chen, Nimish Ronghe et al.· 0 citations
Experiments conducted on movie, book and electronics benchmark datasets demonstrate that GradSup outperforms iterative fine-tuning to provide scalable and personalised recommendation that is consistently sustained above the frozen LLM backbone.
Kanishka Dandeniya, C. Dasanayaka, Daswin de Silva et al.· Proceedings of the 20th ACM...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.