Skip to content

GAUGE: Grading Agent-Built Financial Models Without a Golden Answer

Jul 2026 · arXiv.org · Vol abs/2607.24889 · 1 citation · 34 references
Computer Science

TL;DR

GAUGE is introduced, a benchmark for evaluating agent-built valuation models against observed analyst practice rather than a single point answer, and current agents are substantially stronger at model construction than valuation judgment.

Abstract

Financial models combine public disclosures with analyst assumptions to produce forecasts and valuations. While some components can be checked mechanically, forecasts, discount rates, and target prices often admit multiple reasonable answers. Existing benchmarks nevertheless tend to grade such outputs against a single expert reference. Using independently built analyst models for the same companies, we find that across 108 directed pairs covering 65 companies, the median single-reference score is 0.33, 92.6% score below 0.70, and no same-vintage pair agrees on implied price within 10%. Point-tolerance grading can therefore penalize disagreement already present among professionals. We introduce GAUGE, a benchmark for evaluating agent-built valuation models against observed analyst practice rather than a single point answer. GAUGE uses 1,001 vendor-classified analyst workbooks and a 196-task evaluation set, with a three-layer observed-practice envelope, 56 auditable facets, eight validity gates, and deterministic structural checks. We validate the benchmark with a 55-participant known-groups study, company-grouped cross-fitting, and judge-stability audits. On the failure-aware score $\phi_0$, senior analysts average 88.3, juniors 66.0, and finance students 43.2. Across 24 agents and 1,011 scored generations, the best agent scores 53.4, above the student mean but below every senior and most juniors. It passes 93% of mechanical facets and 78% of judgment facets, with a fleet-median gap of 26 points. Current agents are substantially stronger at model construction than valuation judgment. We release the methodology, a gated de-identified data tier, a controlled training split, a versioned 48-task evaluation core, and a withheld refresh pool.

View source

Similar papers

Preprint Aug 2026

Long-Horizon Forecasting of Complete Financial Statements with Forma

Specialist training beats generalist scale when forecasting financial statements. To our knowledge, no prior work jointly forecasts complete financial statements beyond one year, yet in a discounted-cash-flow valuation most firm value sits past that window. We release ProForma-20Q, a reproducible benchmark for forecast...

Travis L. Johnson, Jian Jiang, Soumyabrata Chaudhuri et al. · 0 citations
Review Aug 2026

The Price of Permission: Classification Uncertainty in Constrained Capital Markets

Shariah-compliant equity screening provides a transparent setting in which institutional rules determine who may own a stock. A binary label identifies current eligibility but not whether the feasible investor base is fragmented across standards or close to changing. We define this instability as classification uncerta...

Abdulrahman Qadi, A. Sharma, Francesca Medda · 0 citations
Preprint Aug 2026

What Makes a Peer? Valuation-Anchored Similarity in Private Markets

An ensemble tree-based supervised similarity learning framework is proposed that defines company similarity through the lens of market valuation rather than static feature matching or semantic descriptions, and improves upon traditional distance-based and text-embedding-based approaches in downstream k-nearest-neighbor...

Sebastian Frank, Jingrao Lyu, Max Jarmey et al. · 0 citations
Preprint Aug 2026

Mandate without Managers: Automated Market Makers as Verifiable Portfolio Products

Automated market makers (AMMs) are typically interpreted and evaluated as decentralized exchanges. Herein, we take the perspective envisioned by Balancer that an AMM can also be viewed as a portfolio technology that programmatically enforces an economic mandate. In particular, we follow the geometric mean market maker...

Zachary Feinstein, I. Florescu, Sean O'Leary · 0 citations
Preprint Aug 2026

FrontierFinance: A Challenging Benchmark for Measuring Frontier Intelligence of Finance Agents

Evaluating FrontierFinance, a fully open benchmark of 220 expert-crafted queries and 11,543 source-attributed rubrics spanning six crucial use cases across the full investor workflow, finds that the tool harness, not the model alone, strongly shapes quality and efficiency.

Yuhao Zhang, O. Koyluoglu, Thejas Venkatesh et al. · 3 citations · ⚡2
Preprint Aug 2026

Calibration-Induced Degeneracy in LLM Financial Forecasting: An Audit-Trailed Case Study on Next-Day Market Risk

Costly LLM features matter only if calibration lets them affect the forecast. We document a failure of this link in a next-day risk study of two broad-market funds. Full-history scoring preceded the 2022 calibration. Calibration then set all four LLM weights to zero. The 856 later scores therefore could not affect the...

A. Mohanty · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.