Inference scaling has been shown to improve large language model (LLM) performance, and this principle naturally extends to autonomous LLM agents through increased search budgets, which we refer to as *search scaling*. Although prior work has characterized the mechanisms, scaling behavior, and performance limits of LLM...
Kang-Cheng Deng, Hui Cai, Jiacheng Lu et al.· 0 citations
CalibratedRubric is introduced, a task-adaptive framework that combines type-specific scoring, Bayesian rubric-measurability filtering, and item response theory (IRT)-based bank assembly that supports CalibratedRubric as an efficient, uncertainty-aware approach to open-ended LLM evaluation.
GAUGE is introduced, a benchmark for evaluating agent-built valuation models against observed analyst practice rather than a single point answer, and current agents are substantially stronger at model construction than valuation judgment.
A scalable pipeline for generating high-quality rubrics without human experts in the final loop is proposed, which is naturally scalable for benchmark evaluation, automatic system comparison, and future studies of evaluation-driven system improvement.
Beidi Luan, Rui Sun, Si-Nuo Wang et al.· arXiv.org· 3 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.