PROMISE: Process Reward Models for Unlocking Test-Time Scaling Laws in Generative Recommendations
This work proposes Promise, a novel framework that integrates dense, step-by-step verification into generative models, and unlocks Test-Time Scaling Laws in recommender systems, demonstrating that by increasing inference compute, smaller models can match or surpass larger models.
Cheng-Cheng Guo, Kuo Cai, Yu Zhou et al.
· Proceedings of the 20th ACM... · 0 citations