RMSWeb: Reflection, Failure-Mode Mining, and Salvage-DS for Web Agent Reinforcement Learning
RMSWeb, a three-part recipe for Qwen3-VL-Instruct at 8B and 32B, achieves the strongest reported Online-Mind2Web result among similarly sized open-weight models in a comparison and a leading reported accuracy-cost trade-off on WebVoyager and WebTailBench, with the caveat that external evaluation protocols differ.