Calibrating the judge is a prerequisite for using citation rubrics as reward signals, and the results show that this calibration does not require the most expensive available model.
Abstract
Reinforcement learning increasingly relies on an LLM judge to score each rubric criterion, and that judge acts as the reward model during training. Before such a signal can be trusted, we need to know how capable the judge must be and how biased it is. We study this calibration question for citation quality in deep-research systems, where a search-grounded LLM must support each claim it writes with a cited source. Citation quality is a structured rubric task in which each attribution-citation pair is judged along two dimensions that require an LLM, source relevance and factual support. On an adversarial long-form benchmark, we score 8 off-the-shelf LLM judges from 3 model families against gold labels over 1,248 rubric decisions, all of which were human-reviewed and 378 of which were hard cases adjudicated from judge disagreements. Cheaper judges remain competitive across both dimensions, with GPT-5-mini attaining the strongest source-relevance pass-class F1 at 0.908 ($\kappa$=0.636), while on factual support the judges are statistically indistinguishable (overlapping confidence intervals), so no single model dominates. At comparable F1, the judges still differ substantially in pass-rate drift, false positive rate, and false negative rate. Scalar F1 obscures this directional bias, yet it is exactly what a downstream reinforcement learning loop would reinforce. Calibrating the judge is therefore a prerequisite for using citation rubrics as reward signals, and our results show that this calibration does not require the most expensive available model.
This work presents HNR-DAC, a two-stage framework that trains each stage on the cases it will actually encounter, and quantifies evidence confusability using a base reranker's scores on non-gold paragraphs and contrasts gold evidence against the most confusable candidates.
Zhenchao Wang, Xin Chen, Luo-Xiu Zhang et al.· 0 citations
This work proposes Rubric Dropout, a one-line fix borrowed from neuron dropout that randomly drops a subset of the rubric's criteria before computing the reward, so the policy never optimizes the same rubric twice.
Minglai Yang, Xinyu Guo, Utkarsh Tyagi et al.· 0 citations
A blockchain-based commit-reveal protocol using Autonomous Economic Agents on an Ethereum-compatible ledger creates a tamper-evident audit trail that separates blind evaluation from post-hoc claims and reduces the verification burden on independent researchers and leaderboard operators.
The Lit2Test benchmark centers on a six-field contract organized around a falsifying outcome, so that every proposal precommits the observation that would prove it wrong, making its quality decidable in the first place rather than merely arguable.
Zi-Yue Wang, Aomufei Yuan, Yi-Ran Yao et al.· 2 citations
HALLMARK (Hallucination benchmark): 2,526 BibTeX entries spanning 14 hallucination types, three difficulty tiers, six diagnostic sub-tests per entry, and a contamination-resistant held-out split, and across the benchmark one result is consistent: the false-positive rate, not recall, decides whether a verifier is deploy...
This work studies whether smaller language models can serve as efficient and reliable rubric-based judges, and compares three ways of extracting criterion-level judgments from small models: Generative verdicts, Yes/No Logprob margins, and Probe judges.
Feng-Yu Xie, Yilun Zhao, Bingsen Chen et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.