Skip to content

CalibratedRubric: Task-Adaptive Rubric Banks for Open-Ended LLM Evaluation

Jul 2026 · arXiv.org · Vol abs/2607.29252 · 0 citations
Computer Science

TL;DR

CalibratedRubric is introduced, a task-adaptive framework that combines type-specific scoring, Bayesian rubric-measurability filtering, and item response theory (IRT)-based bank assembly that supports CalibratedRubric as an efficient, uncertainty-aware approach to open-ended LLM evaluation.

Abstract

Reliable evaluation of open-ended LLM outputs requires fine-grained rubrics, yet expert curation is costly and difficult to scale. Existing automated pipelines rely on strict judge unanimity and binary variance filters, which cannot distinguish measurable rubrics from informative ones. We introduce CalibratedRubric, a task-adaptive framework that combines type-specific scoring, Bayesian rubric-measurability filtering, and item response theory (IRT)-based bank assembly. CalibratedRubric estimates each rubric's measurability with a Beta--Bernoulli agreement posterior and uses a submodular information-coverage objective to construct compact rubric banks over the observed capability range. Across financial, healthcare, general, and legal benchmarks, measurability filtering improves human-gold agreement on JudgmentBench from $\kappa=0.604$ to $0.743$. IRT-based greedy selection improves cross-fitted rank fidelity over random selection across all six evaluated response blocks and requires only 49 rather than 131 rubrics to reach the target correlation on FinResearchBench decision-support tasks. Task-label perturbations further reduce system separation, confirming the practical relevance of task-adaptive scoring. These results support CalibratedRubric as an efficient, uncertainty-aware approach to open-ended LLM evaluation, with calibration gains depending on sufficient judge redundancy.

View source

Similar papers

Open access Aug 2026

CreditTrace-LLM: Auditing Rubric-Point Responsiveness and Credit Locality in LLM-Based Automated Grading

CreditTrace-LLM evaluates target responsiveness, credit locality, paraphrase stability, and traceable Gold-point-level outputs as complementary diagnostic dimensions rather than as predefined criteria for acceptable grading performance.

Cătălin Anghel, A. Anghel, M. Craciun et al. · 0 citations
#natural language process... Preprint Aug 2026

GenRubric: Self-Evolving Rubric Generation for Scalable LLM Evaluation

GenRubric is introduced, a self-evolving framework that improves rubric generation from unlabeled queries without requiring additional human annotations during self-evolution, and experiments show that self-evolution improves the agreement between evaluations induced by generated rubrics and those induced by expert-wri...

Yifan Chen, Hai-Tao Li, Qing-Yao Ai et al. · 2 citations · ⚡1
#artificial intelligence Preprint Sep 2026

CARE: Contrastive Anchor-based Rubric Evolution for Large Language Model Post-Training

CARE is proposed, which grounds every rubric evolution step in a high-quality anchor response generated by a frontier model conditioned on the prompt and its rubrics, enabling two complementary mechanisms: an Adaptive branch that reactively repairs reward misspecification; and a Chase branch that proactively converts f...

Si-Yuan Li, Xin-Xin Song, Rui-Nian Chen et al. · 0 citations
Review Aug 2026

JudgeStealer: Extracting LLM Judging Capabilities across Evaluation Protocols

JUDGESTEALER is proposed, the first query-efficient model extraction framework for replicating judging capabilities across pointwise scoring, pairwise comparison, and listwise ranking protocols and demonstrates robustness against representative extraction defenses.

Chen Chen, Yao-Lin Chen, Xue-Han Sun et al. · 0 citations

Clustering-based Prompt Optimization for LLM Evaluation

The cluster loop yields the strongest held-out rubric on both evaluator tasks from a commercial search vertical, and is the only method robustly positive on both.

Jinyoung Kim, N. Corp, Sun Kim et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.