Skip to content
Preprint

RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning

Aug 2026 · 0 citations · 36 references
Computer Science

TL;DR

Results indicate that selective internalization through reward filtering and policy support shaping is effective for open-ended reinforcement learning.

Abstract

Aligning Large Language Models (LLMs) for open-ended tasks is challenging because responses must satisfy multidimensional criteria without following a single correct generation trajectory. Existing rubric-based reinforcement learning (RL) methods compress fine-grained criterion-level feedback into scalar rewards, making persistent capability gaps difficult to target under limited on-policy exploration. We propose $\textbf{RISE-RL}$ (Rubric-Informed Selective Exploration), which uses repeatedly missed rubric criteria to elicit privileged trajectories that are difficult to discover through unguided exploration alone. RISE-RL retains only trajectories whose complete-rubric reward exceeds the mean reward of natural rollouts, and then re-evaluates them under the original prompt to emphasize behaviors that remain weakly supported by the natural policy. The resulting guidance signal is optimized through a separate auxiliary objective and removed once its additional benefit diminishes. Experiments with 4B and 14B models across writing, chat, health, and science show that RISE-RL achieves the highest mean score on every evaluated benchmark under guidance-free evaluation. Compared with standard Rubric-RL, it improves the average score by 1.3 points at the 4B scale and $\textbf{3.3 points at the 14B scale}$, including a $\textbf{6.0-point}$ gain on CreativeWriting-V3. It also improves creative-writing diversity and yields gains on objectively scored medical and scientific benchmarks. These results indicate that selective internalization through reward filtering and policy support shaping is effective for open-ended reinforcement learning.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

ARISE-RL: Agentic Rubric-Grounded Iterative Self-Evolution with Reinforcement Learning

ARISE-RL, a novel full-cycle self-evolution framework that couples a task/rubric Generator and a reasoning Solver through rubric-mediated co-evolution, is proposed and ECR-Bench, an expert-calibrated rubric benchmark suite covering single-tool deep research and multi-tool travel planning is presented.

Fanrui Zhang, Rui-Xue Ding, Qiang Zhang et al. · 0 citations
#artificial intelligence Review Aug 2026

A Survey on Rubric-Guided Reinforcement Learning for Language Models

A Bayesian framework that defines constitutions as prior distributions over evaluation criteria and rubrics as conditional instantiations is introduced, and a taxonomy of rubric-guided RL along the prior-posterior axis is presented, covering constitutional AI, instance-specific rubrics, process-level supervision, self-...

Zifei Shan, Fang-Ning Shao · 0 citations
Jul 2026

SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning

Test-time reinforcement learning (TTRL) enables language models to self-evolve at inference time without labeled feedback. Existing methods rely on answer voting and therefore do not extend naturally to open-ended generation, where valid responses cannot be mapped to a shared canonical answer. Without external reward m...

Jianze Wang, Kun-Wang Zheng, Ying Liu et al. · 0 citations
Preprint Aug 2026

Rubrics as Privileged Information for Open-Ended Generation

It is shown that soft rubric PI provides a larger and more effective training signal on student roll-outs than hard reference completion PI in this regime, and contrary to intuition, soft rubric PI provides a larger and more effective training signal than hard reference completion PI in this regime.

Deepika Bablani, Ajay Gupta, Wan-Xuan Chen · 1 citation
Jul 2026

H2SD: Hybrid Hindsight Self-Distillation

Experiments on challenging reasoning benchmarks show that H$^2$SD achieves the strongest overall performance among representative RLVR and self-distillation baselines, with stable optimization and a favorable accuracy-efficiency trade-off.

Qi Cai, Yi-Chuan Ma, Linyang Li et al. · 2 citations
#artificial intelligence Preprint Sep 2026

Environments as Scaffold: Enriching Feedback to Bootstrap Self-Evolving Agents in Long-Horizon Tasks

Large Language Models demonstrate remarkable proficiency in static reasoning, yet training them as autonomous agents through Reinforcement Learning (RL) for long-horizon tasks is often hindered by severe reward sparsity. While conventional \textit{agent-side warming} up via supervised fine-tuning (SFT) can alleviate th...

Hongbang Yuan, Zhuo-Ran Jin, Yi-Xin Cao · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.