Generative recommendation, has been attracted a surge of attentions in industrial and academic research community, towards to build more smart system to build next-generation recommender. Under the significant developing wave of large language model, our team have been developed Semantic ID based OneRec/OneRec-V2. Thes...
Jiang-Xia Cao, Hao Peng, Wen-Long Xu et al.· 0 citations
These results show that RecoReward trains the MLLM to produce item features that benefit downstream recommendation while retaining content-only serving, and shows that RecoReward-9B outperforms its Qwen3.5-9B baseline and all other evaluated models across seven recall metrics.
Guohong Mu, Yue-Yang Liu, Jiangxia Cao et al.· arXiv.org· 0 citations
The sample efficiency and scalability of RL post-training for video MLLMs and introduces OraRL, a decoupled advantage estimator that scales with model size and data, surpassing its backbone from 0.8B to 9B and GRPO up to 100k prompts.
Yunheng Li, Guo-Hong Mu, Hao Li et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.