RMSWeb, a three-part recipe for Qwen3-VL-Instruct at 8B and 32B, achieves the strongest reported Online-Mind2Web result among similarly sized open-weight models in a comparison and a leading reported accuracy-cost trade-off on WebVoyager and WebTailBench, with the caveat that external evaluation protocols differ.
Abstract
Compact web agents can reduce deployment cost, but training them poses challenges in both data collection and post-SFT reinforcement learning (RL). Successful trajectories are expensive to collect and often contain inefficient detours. After supervised fine-tuning (SFT), full trajectory corpora are dominated by routine states; moreover, when group-relative RL is applied to web actions, inadequately designed action-level rewards can yield weak or misleading relative updates, while groups rejected as unsuitable for such updates receive no fallback learning signal. We present RMSWeb, a three-part recipe for Qwen3-VL-Instruct at 8B and 32B. Reflection-conditioned retries increase collection yield and shorten successful trajectories; failure-mode mining concentrates offline RL on critical states exposed by the SFT policy; and Salvage-DS combines an action-semantic polarized reward, contrast-and-competence-gated dynamic sampling, and an action-only anchor for rejected groups. Policies trained with reflection-collected data use up to 19.7% fewer action steps on solved tasks. On WebVoyager, Online-Mind2Web, and WebTailBench, RMSWeb improves over SFT by 2.4-7.0 points at 8B and 1.2-7.7 points at 32B. Our 8B model also achieves the strongest reported Online-Mind2Web result among similarly sized open-weight models in our comparison and a leading reported accuracy-cost trade-off on WebVoyager and WebTailBench, with the caveat that external evaluation protocols differ.
Self-improving federated agent networks keep training after deployment by collecting new trajectories with the current policy and feeding them back into later rounds. This closed loop makes unlearning harder than a one-time model repair. When a data owner requests deletion, the target data may have already shaped later...
Zihao Ding, Jun Huang, Liang Dong· arXiv.org· 0 citations
T1, a Mixture-of-Experts model of 122B total trained with reinforcement learning, operating a real shell in a cloud sandbox for up to 300+ tool-call turns per task, rewarded by executing each task's own verifier by executing each task's own verifier.
Jun-Yao Yang, Yu-Cheng Shi, Zhong-Zhi Li et al.· 1 citation
CANOPY (Coverage-ANchored On-PolicY RL), a minimalist protocol attacking both Signal starvation and policy drift, internalizes long-horizon capability directly into a small open model; the complete training stack is planned to be released at https://github.com/AlibabaResearch/SignalCoverageRL.
Li-Ming Pu, Xiao-Xiao Li, Yi-Fu Liu et al.· 0 citations
I-SDPO (Instance-Level Adaptive Self-Distillation Policy Optimization), which treats teacher reliance as capability-dependent and uses imitation only where group-relative rewards are uninformative, obtains the best result in all four scientific domains.
Yubo Zhang, Xin-Hong Ma, Zezhong Tan et al.· 0 citations
This work casts per-step rollout collection as a budget-constrained sequential allocation problem and introduces SARA (Sequential Adaptive Rollout Allocation), a two-threshold, SPRT-style rule that commits effective groups, abandons saturated ones after a short probe, and reallocates the freed budget to fresh prompts,...
Pixel Nomand, Elena Voss, Marcus Hale et al.· arXiv.org· 0 citations
DIEM is proposed, a principled and fully automated framework that makes data utilization adaptive throughout RFT and consistently outperforms strong static and dynamic baselines.
Haoru Tan, Sitong Wu, Yan-Feng Chen et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.