Inference scaling has been shown to improve large language model (LLM) performance, and this principle naturally extends to autonomous LLM agents through increased search budgets, which we refer to as *search scaling*. Although prior work has characterized the mechanisms, scaling behavior, and performance limits of LLM...
Kang-Cheng Deng, Hui Cai, Jiacheng Lu et al.· 0 citations
While Large Language Models (LLMs) achieve high accuracy on established Classical Chinese Poetry benchmarks, it remains challenging to distinguish transferable Linguistic-Aesthetic Reasoning from reliance on familiar pre-training patterns. To address this issue, we introduce Neo-Classic, an evaluation benchmark that co...
Han Zhang, Zihan Gu, Zhiyuan Wang et al.· Annual Meeting of the Associ...· 0 citations
GAUGE is introduced, a benchmark for evaluating agent-built valuation models against observed analyst practice rather than a single point answer, and current agents are substantially stronger at model construction than valuation judgment.