FinIndices is a large-scale benchmark evaluating data-processing fidelity over uncropped financial statements (up to 32K tokens) and yields substantial zero-hint gains, validating that structured logic can be partially restored via data-centric alignment.
Results support treating PHD as a distinct target for evaluating how LLMs formulate investigative directions when final conclusions are withheld, with gains for several lower-performing models on HypoArena but regressions for other systems, including a top-performing model.
Tianyun Zhong, Wangyi Jiang, Wei Wang et al.· arXiv.org· 0 citations
A diagnostic benchmark for open-domain, open-form, long-horizon counterfactual causal reasoning, containing 220 what-if questions across STEM, HSS, and Hybrid scenarios, and finds that WhatIfBench remains far from saturated: even the strongest model reaches only a 64.62% final score.
Yucheng Wang, Yuetian Du, Zheng Liu et al.· 0 citations
Living-Harness is proposed, a self-evolving agent harness that converts each completed trajectory and its evaluator signals into posterior evidence for bounded harness updates, and supports retrieval-only reuse of the evolved harness state across model backbones.
Yuetian Du, Yucheng Wang, Helsing Xu et al.· arXiv.org· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.