Skip to content
Book Open access

TransAct V2: Production System for Lifelong User Sequence Modeling at Scale

Sep 2026 · Proceedings of the 20th ACM Conference on Recommender Systems · 0 citations · 17 references

TL;DR

This work presents TransAct V2, a production ranking model deployed at Pinterest serving over 630 million users, that addresses challenges through a dual-path training-serving architecture with request-level deduplication and int8 quantization and provides a reproducible blueprint for practitioners building large-scale sequential recommenders.

Abstract

Modeling lifelong user action sequences in CTR prediction faces three production challenges: computational complexity with quadratic transformer costs as sequences extend from hundreds to tens of thousands of actions, infrastructure overhead from storage and network costs that scale with both sequence length and number of candidates per request, and weak training supervision as sequence encoders sit far from CTR prediction heads. We present TransAct V2, a production ranking model deployed at Pinterest serving over 630 million users, that addresses these challenges through: (1) a dual-path training-serving architecture with request-level deduplication and int8 quantization achieving 1% logging cost, (2) fused Triton kernels with pinned memory management delivering 250 × p99 latency improvement, and (3) a Next Action Loss providing direct supervision for sequence modeling within CTR frameworks. Deployed in production, TransAct V2 demonstrates significant improvements in both engagement and recommendation diversity. To support industry adoption, we open-source our serving optimizations with comprehensive ablation studies. Our work provides a reproducible blueprint for practitioners building large-scale sequential recommenders.

Read PDF

Similar papers

#artificial intelligence Preprint Sep 2026

From Language to Behavior: Scaling Sequence Transformers for Industrial Recommendation Ranking with Rec-Native Designs

ReST is proposed, a recommendation-native Transformer scaling framework that achieves higher accuracy and scales more consistently along sequence length, depth, and width, where LLM-style Transformer blocks saturate.

Jie Chen, Xiang-Qia Yu, Yan-Chao Lian et al. · 0 citations
Book Open access Sep 2026

UniTraj: Cross-Domain Long-Sequence Modeling for Commercial Recommendation

UniTraj, a practical framework that extends sequence construction beyond the advertising domain by incorporating behaviors from content-consumption scenarios, forming unified commercial trajectories across domains and scenarios, is proposed and deployed in a large-scale online advertising system.

Xian Hu, Ming Yue, Zhi-Xiang Feng et al. · 0 citations
#machine learning Preprint Sep 2026

KuaFu: Compressing Long User Behavior into Understanding at Billion Scale

KuaFu is a unified behavior-compression layer whose minimal unit is one behavior item, with fidelity-oriented four-stage training and layered intermediate evaluation, which matches or exceeds uncompressed single-task production models on all five headline metrics, raises per-GPU throughput by 37%-350%, and saves 190 GP...

Jia-Hao Hui, Lin Zhu, Yi-Sheng Hu et al. · 0 citations
Book Open access Sep 2026

Interest Sequence for User Modeling in Industrial Short-Form Video Recommendation

Ultra-long user history modeling has been a highly effective approach in modern industrial recommendation systems, with most works heavily utilizing search-based methods and summarization based methods to map massive user interaction logs into latent user representation with algorithm designs to address the scaling cha...

Yuan-Zheng Lin, Diego Uribe Mora, Yuan Shao et al. · 0 citations
#machine learning Preprint Aug 2026

SAGA: Structure-Attended Generative Action Embedding Model that encodes Multi-Surface User Action Sequences

This work presents SAGA, a generative action embedding model that encodes multi-surface user interaction sequences across a Financial Service organization's ecosystems, from checkout, peer-to-peer (P2P) transactions, in-app engagement, email to account actions, into a unified user representation for downstream recommen...

Tsz-Fung Pang, Po‐Jui Chen, Nimish Ronghe et al. · 0 citations
Book Open access Sep 2026

GradSup: Gradient Superposition for Personalised and Scalable LLM Recommendation

Experiments conducted on movie, book and electronics benchmark datasets demonstrate that GradSup outperforms iterative fine-tuning to provide scalable and personalised recommendation that is consistently sustained above the frozen LLM backbone.

Kanishka Dandeniya, C. Dasanayaka, Daswin de Silva et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.