A tailored Bootstrapped Pareto Policy Optimization (BPPO) is proposed, which synergizes Bootstrapping Reward Rectification and Conflict-Aware Pareto Advantage Fusion (CPAF) and exposes critical reasoning deficits, highlighting imperative for knowledge-intensive benchmarks towards next-generation visual generation.
Yuan Wang, Yongchao Du, Mengting Chen et al.· arXiv.org· 1 citation
This work introduces MMShopBench, the first real-log benchmark for multimodal, multi-turn shopping agents, and builds an offline shopping sandbox, where fine-tuning substantially narrows the performance gap between the open-source model and leading proprietary models.
This work proposes CMI-Mem, a lightweight RL memory manager with a hybrid reward, which demonstrates improved transfer across memory-use scenarios, together with more efficient training and inference from the per-operation CMI signal.
Yubo Wang, Qiuyu Zhao, Zenghui Sun et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.