CalibratedRubric is introduced, a task-adaptive framework that combines type-specific scoring, Bayesian rubric-measurability filtering, and item response theory (IRT)-based bank assembly that supports CalibratedRubric as an efficient, uncertainty-aware approach to open-ended LLM evaluation.
Abstract
Reliable evaluation of open-ended LLM outputs requires fine-grained rubrics, yet expert curation is costly and difficult to scale. Existing automated pipelines rely on strict judge unanimity and binary variance filters, which cannot distinguish measurable rubrics from informative ones. We introduce CalibratedRubric, a task-adaptive framework that combines type-specific scoring, Bayesian rubric-measurability filtering, and item response theory (IRT)-based bank assembly. CalibratedRubric estimates each rubric's measurability with a Beta--Bernoulli agreement posterior and uses a submodular information-coverage objective to construct compact rubric banks over the observed capability range. Across financial, healthcare, general, and legal benchmarks, measurability filtering improves human-gold agreement on JudgmentBench from $\kappa=0.604$ to $0.743$. IRT-based greedy selection improves cross-fitted rank fidelity over random selection across all six evaluated response blocks and requires only 49 rather than 131 rubrics to reach the target correlation on FinResearchBench decision-support tasks. Task-label perturbations further reduce system separation, confirming the practical relevance of task-adaptive scoring. These results support CalibratedRubric as an efficient, uncertainty-aware approach to open-ended LLM evaluation, with calibration gains depending on sufficient judge redundancy.
CreditTrace-LLM evaluates target responsiveness, credit locality, paraphrase stability, and traceable Gold-point-level outputs as complementary diagnostic dimensions rather than as predefined criteria for acceptable grading performance.
Cătălin Anghel, A. Anghel, M. Craciun et al.· Computers· 0 citations
GenRubric is introduced, a self-evolving framework that improves rubric generation from unlabeled queries without requiring additional human annotations during self-evolution, and experiments show that self-evolution improves the agreement between evaluations induced by generated rubrics and those induced by expert-wri...
Yifan Chen, Hai-Tao Li, Qing-Yao Ai et al.· 2 citations· ⚡1
CARE is proposed, which grounds every rubric evolution step in a high-quality anchor response generated by a frontier model conditioned on the prompt and its rubrics, enabling two complementary mechanisms: an Adaptive branch that reactively repairs reward misspecification; and a Chase branch that proactively converts f...
Si-Yuan Li, Xin-Xin Song, Rui-Nian Chen et al.· 0 citations
JUDGESTEALER is proposed, the first query-efficient model extraction framework for replicating judging capabilities across pointwise scoring, pairwise comparison, and listwise ranking protocols and demonstrates robustness against representative extraction defenses.
Chen Chen, Yao-Lin Chen, Xue-Han Sun et al.· 0 citations
NovGauge is presented, a human-anchored benchmark for fine-grained novelty assessment diagnosis, and a cascading diagnostic pipeline that verifies per-dimension correctness, evidence grounding, and logical support is proposed.
Guo-Qiang Zhang, Ke-Xin Tan, Ming Zhang et al.· 0 citations
The cluster loop yields the strongest held-out rubric on both evaluator tasks from a commercial search vertical, and is the only method robustly positive on both.
Jinyoung Kim, N. Corp, Sun Kim et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.