A scalable pipeline for generating high-quality rubrics without human experts in the final loop is proposed, which is naturally scalable for benchmark evaluation, automatic system comparison, and future studies of evaluation-driven system improvement.
Abstract
Deep research agents are increasingly used to produce long-form financial reports, yet large-scale evaluation remains bottlenecked by the need for human experts to define and execute high-quality rubrics. We address this problem by proposing a scalable pipeline for generating high-quality rubrics without human experts in the final loop. We build a financial deep research benchmark from 104 real-world user queries and automatically synthesize 14,450 query-specific candidate rubrics from model-generated reports. To justify removing human experts from rubric execution, we compare rubric judgments from three human experts with those from a three-LLM judge panel on a sampled subset, and show that LLM-based evaluation is sufficiently consistent with human evaluation to replace it for large-scale rubric screening, including 98.67\% label-level agreement on jointly unanimous items. We then derive consensus-derived gold rubrics through two filters: a strict consistency filter, which keeps a rubric only if the three LLM judges unanimously agree on every report under the same query, and a distinguishability filter, which keeps a rubric only if it assigns at least one majority-yes and at least one majority-no label across the evaluated systems. This process retains 3,687 consistency-passed rubrics, of which 2,600 remain distinguishable and form the final set of consensus-derived gold rubrics. Using this final rubric set, we obtain clearly differentiated rankings across 10 deep research systems, with item-level pass rates ranging from 58.58\% to 22.23\%. More broadly, because the pipeline removes human-expert execution from rubric generation and evaluation, it is naturally scalable for benchmark evaluation, automatic system comparison, and future studies of evaluation-driven system improvement.
GenRubric is introduced, a self-evolving framework that improves rubric generation from unlabeled queries without requiring additional human annotations during self-evolution, and experiments show that self-evolution improves the agreement between evaluations induced by generated rubrics and those induced by expert-wri...
Yifan Chen, Hai-Tao Li, Qing-Yao Ai et al.· 2 citations· ⚡1
This work introduces a verifiable benchmark of 500 deep research tasks spanning 31 topics and 10 major categories, with three query forms designed to probe complementary capabilities required for deep research.
Can Wang, Hao-Ran Chen, Hao Gao et al.· 0 citations
The cluster loop yields the strongest held-out rubric on both evaluator tasks from a commercial search vertical, and is the only method robustly positive on both.
Jinyoung Kim, N. Corp, Sun Kim et al.· 0 citations
This work studies whether smaller language models can serve as efficient and reliable rubric-based judges, and compares three ways of extracting criterion-level judgments from small models: Generative verdicts, Yes/No Logprob margins, and Probe judges.
Feng-Yu Xie, Yilun Zhao, Bingsen Chen et al.· 0 citations
CalibratedRubric is introduced, a task-adaptive framework that combines type-specific scoring, Bayesian rubric-measurability filtering, and item response theory (IRT)-based bank assembly that supports CalibratedRubric as an efficient, uncertainty-aware approach to open-ended LLM evaluation.
WANDR (Wide ANd Deep Research) is a benchmark of 500 realistic, challenging data-collection tasks for research agents that replaces static gold answer sets with task-specific judges that refetch cited pages and verify each record against its evidence, allowing evaluation of current and changing facts.
Vitaliy Polshkov, Marcin Pitera, Jeremy Yang et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.