Skip to content
Preprint

Rubrics as Privileged Information for Open-Ended Generation

Aug 2026 · 1 citation · 19 references
Computer Science

TL;DR

It is shown that soft rubric PI provides a larger and more effective training signal on student roll-outs than hard reference completion PI in this regime, and contrary to intuition, soft rubric PI provides a larger and more effective training signal than hard reference completion PI in this regime.

Abstract

On-policy self-distillation (OPSD), where a single model acts as both student and teacher with different contexts, has shown promise in verifiable domains like math, where hard privileged information (PI) in the form of ground-truth answers structurally constrains valid continuations. We extend OPSD to open-ended generation using soft PI in the form of rubrics that guide preferences but admit many valid responses. Rubrics have served as scalar rewards for reinforcement learning (RL); we show that they provide substantially richer signal as dense PI for distillation, and contrary to intuition, soft rubric PI provides a larger and more effective training signal on student roll-outs than hard reference completion PI in this regime. A reference completion is one point in a set of valid responses, so distilling towards it over-constrains the student, while rubrics specify the preference structure shared across the set of valid responses. We show the effectiveness of using rubrics as PI for open-ended generation across Qwen and Llama model families and show that it outperforms rubric-as-reward (RaR) RL using HealthBench, a benchmark that grades open-ended health responses against physician-created rubrics, providing dense token-level supervision for open-ended tasks; RuPI beats RaR RL by up to +0.10 absolute score and, under matched recipe and KL direction, beats reference-PI by +0.034 to +0.079 absolute score across three models. We further show that these findings generalize to training on the RubricHub Science corpus and evaluating on ResearchQA: soft rubric PI outperforms both reference-PI distillation and RaR RL (66.6% vs. 64.2% and 57.6%).

View source

Similar papers

Preprint Aug 2026

Privileged, but Biased: How PI-Conditioned Teachers Break Self-Distillation

This work reproduces SDPO's reported gains in its easy setting, then applies the identical setup to difficult tasks and finds that it does not teach anything, and explains this failure through a single causal chain from the loss to the model it produces.

Sarthak Harne, Chinmay Karkar, Yash Pandya et al. · 6 citations
Jul 2026

LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks

Experiential Learning is proposed, which repurposes the feedback model from an LLM-as-a-Judge into an LLM-as-a-Coach, and establishes experiential knowledge as a richer and more generalizable learning signal for post-training on non-verifiable tasks.

Tianzhu Ye, Li Dong, Guanheng Chen et al. · 1 citation
#machine learning Preprint Sep 2026

Negative Self-Distillation: Learning to Reason by Avoiding Flaws

On-Policy Self-Distillation (OPSD) has emerged as a popular paradigm for large language model (LLM) self-improvement, allowing models to act as their own teachers by leveraging privileged information such as ground-truth solutions. However, recent findings indicate that OPSD can severely degrade the performance of LLMs...

Rong-Can Pei, Zhepei Wei, Shu-Yao Xu et al. · 1 citation
Jul 2026

H2SD: Hybrid Hindsight Self-Distillation

Experiments on challenging reasoning benchmarks show that H$^2$SD achieves the strongest overall performance among representative RLVR and self-distillation baselines, with stable optimization and a favorable accuracy-efficiency trade-off.

Qi Cai, Yi-Chuan Ma, Linyang Li et al. · 2 citations
#artificial intelligence Preprint Sep 2026

DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics for Long-Horizon Agent Training

Reinforcement Learning from Verifiable Rewards works well when a task has a programmatic checker, but most long-horizon agent domains have none. We work in the outcome-blind setting, where ground-truth success signals are not available. Multi-criteria rubrics are a popular way to supply such a reward; they are scored o...

Shubham Gandhi, Saurabh Goyal, K. Kate et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.