Skip to content
Review

Do You Need a Frontier Model as a Citation Verifier? Benchmarking Rubric LLMs for Deep-Research Source Attribution

Jul 2026 · arXiv.org · Vol abs/2607.08700 · 0 citations · 34 references
Computer Science

TL;DR

Calibrating the judge is a prerequisite for using citation rubrics as reward signals, and the results show that this calibration does not require the most expensive available model.

Abstract

Reinforcement learning increasingly relies on an LLM judge to score each rubric criterion, and that judge acts as the reward model during training. Before such a signal can be trusted, we need to know how capable the judge must be and how biased it is. We study this calibration question for citation quality in deep-research systems, where a search-grounded LLM must support each claim it writes with a cited source. Citation quality is a structured rubric task in which each attribution-citation pair is judged along two dimensions that require an LLM, source relevance and factual support. On an adversarial long-form benchmark, we score 8 off-the-shelf LLM judges from 3 model families against gold labels over 1,248 rubric decisions, all of which were human-reviewed and 378 of which were hard cases adjudicated from judge disagreements. Cheaper judges remain competitive across both dimensions, with GPT-5-mini attaining the strongest source-relevance pass-class F1 at 0.908 ($\kappa$=0.636), while on factual support the judges are statistically indistinguishable (overlapping confidence intervals), so no single model dominates. At comparable F1, the judges still differ substantially in pass-rate drift, false positive rate, and false negative rate. Scalar F1 obscures this directional bias, yet it is exactly what a downstream reinforcement learning loop would reinforce. Calibrating the judge is therefore a prerequisite for using citation rubrics as reward signals, and our results show that this calibration does not require the most expensive available model.

View source

Similar papers

Preprint Aug 2026

HNR-DAC: Hard-Negative Reranking and Distribution-Aligned Classification for Scientific Claim Verification

This work presents HNR-DAC, a two-stage framework that trains each stage on the cases it will actually encounter, and quantifies evidence confusability using a base reranker's scores on non-gold paragraphs and contrasts gold evidence against the most confusable candidates.

Zhenchao Wang, Xin Chen, Luo-Xiu Zhang et al. · 0 citations
Preprint Aug 2026

Rubric Dropout: A Simple Way to Mitigate Reward Hacking in Rubric-as-Reward RL

This work proposes Rubric Dropout, a one-line fix borrowed from neuron dropout that randomly drops a subset of the rubric's criteria before computing the reward, so the policy never optimizes the same rubric twice.

Minglai Yang, Xinyu Guo, Utkarsh Tyagi et al. · 0 citations
Review Aug 2026

Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation

A blockchain-based commit-reveal protocol using Autonomous Economic Agents on an Ethereum-compatible ledger creates a tamper-evident audit trail that separates blind evaluation from post-hoc claims and reduces the verification burden on independent researchers and leaderboard operators.

Sahil Pardasani, Madhusudan Singh · 0 citations
Preprint Aug 2026

What Proves You Wrong: Benchmarking Language Models on Falsifiable Research Ideation

The Lit2Test benchmark centers on a six-field contract organized around a falsifying outcome, so that every proposal precommits the observation that would prove it wrong, making its quality decidable in the first place rather than merely arguable.

Zi-Yue Wang, Aomufei Yuan, Yi-Ran Yao et al. · 2 citations
Review Jul 2026

HALLMARK: Diagnosing Three Failure Modes in LLM Citation Verifiers

HALLMARK (Hallucination benchmark): 2,526 BibTeX entries spanning 14 hallucination types, three difficulty tiers, six diagnostic sub-tests per entry, and a contamination-resistant held-out split, and across the benchmark one result is consistent: the false-positive rate, not recall, decides whether a verifier is deploy...

Patrik Reizinger, Wieland Brendel · 1 citation
#natural language process... Preprint Aug 2026

Small Language Models as Judges for Rubric-Based Reinforcement Learning

This work studies whether smaller language models can serve as efficient and reliable rubric-based judges, and compares three ways of extracting criterion-level judgments from small models: Generative verdicts, Yes/No Logprob margins, and Probe judges.

Feng-Yu Xie, Yilun Zhao, Bingsen Chen et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.