Skip to content
Preprint

Improving Generalization Robustness of Multimodal RLVR

Aug 2026 · 0 citations · 39 references
Computer Science

Abstract

Reinforcement Learning with Verifiable Rewards (RLVR) makes Multimodal Large Language Models more accurate, but the gains are brittle: simply paraphrasing a question or changing the prompt template can degrade them, which challenges reliable deployment in high-stakes scenarios like medical VQA. We trace this to two issues of the standard RL objective. First, the binary verifier conflates format with content, so the reward signal cannot tell a wrong answer apart from a misformatted one. Second, the training distribution covers only a thin slice of the real-world prompts that the model might meet at deployment, so policies that perform well on the training distribution can behave differently under unseen prompts during test. Both failures call for a robust post-training method that helps the policy cover a broader distribution of semantically equivalent prompts, and we identify two measures that help achieve this objective: separating format from semantics in the reward, and applying policy invariance across perturbed prompts with equivalent semantics. We therefore propose Prompt-Invariant RLVR (PIRL), consisting of a dynamic trinary reward and a consistency regularizer based on an embedding-space adversary. Under stress testing, PIRL's average accuracy on benchmarks drops by only $\le 1\%$, where GRPO drops ~3%. On dynamic evaluation, PIRL also achieves the smallest performance drop.

View source

Similar papers

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information

Off-Context GRPO (OC-GRPO), a minimally modified variant of GRPO that uses guided rollouts but applies an importance-corrected objective to steer the update back toward the original unguided objective, avoiding the mismatch that destabilizes uncorrected guided training.

Priyank Agrawal, Ankur Samanta, S. Ghasemlou et al. · 1 citation
#artificial intelligence Preprint Aug 2026

On-policy Distillation with Verifiable Reward

This work reformulates the implicit reward of sampled-token OPD based on trajectory correctness, then applies a ReLU gating mechanism to ensure that correct trajectories receive non-negative rewards and incorrect ones receive non-positive rewards, making it readily combinable with any policy gradient algorithm, such as...

Wen-Ze Lin, Jia-Le Zhao, Xi-Tai Jiang et al. · 9 citations · ⚡4
Preprint Aug 2026

Annotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMs

The sample efficiency and scalability of RL post-training for video MLLMs and introduces OraRL, a decoupled advantage estimator that scales with model size and data, surpassing its backbone from 0.8B to 9B and GRPO up to 100k prompts.

Yunheng Li, Guo-Hong Mu, Hao Li et al. · 1 citation
#natural language process... Preprint Aug 2026

Consolidating RLVR Capabilities Across Domains: A Deep Dive into Fusion Paradigms

This work compares three fusion paradigms by the artefacts they reuse and suggests that Merge when experts already exist and cheap fusion is paramount; Mix RL when training a unified model without experts, with domain proportions adjusted for cross-domain transfer; and MOPD when preserving domain-specific gains matters...

Sicheng Wu, Kai Yang, Yuchen Cai et al. · 0 citations
#small language model Preprint Aug 2026

Boosting LLM Exploration via Weak-Model Guidance in RLVR

This work empirically study the potential of outer prefixes, revealing the mechanism of the impact of distributional discrepancy to the exploration dynamics in RLVR training and efficiently mitigates entropy collapse without requiring additional SFT, intricate reward designs, or complex prompting.

Xin Shen, Huishuai Zhang, Peng Li et al. · 1 citation
#machine learning Preprint Sep 2026

Cliff: Learning Process Rewards from the First Mistake

Cliff, a reward shaping strategy that utilizes an off-the-shelf LLM as a teacher to identify the first mistake in each rollout, is proposed and established as a simple, general and effective approach for improving RLVR with richer, fine-grained supervision.

Pei-Xuan Han, Runnan Wang, Ketan Ramaneti et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.