Modern recommender systems advance not only by scaling data and parameters, but also by encoding task-specific inductive biases through architecture, including sparse feature interactions for click-through rate (CTR) prediction, temporal attention for sequential recommendation, and expert routing for multi-task learnin...
Xiao-Peng Li, Kuo Cai, Bo Chen et al.· 0 citations
Scaling model capacity has emerged as an effective approach to overcoming performance bottlenecks in industrial recommender systems. However, repeatedly training larger dense models from scratch demands substantial data and time, while their growing computation conflicts with the strict serving budgets of industrial sy...
Rui-Hao Zhang, Bo Chen, Xiao Wang et al.· 0 citations
A lightweight training framework that learns a single Behavior-Equivalent Token that substantially reduces inference cost and frees nearly the entire context window for user inputs and model outputs.
Jiancheng Dong, Pengyue Jia, Jingyu Peng et al.· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.