Skip to content
Preprint

FrontierFinance: A Challenging Benchmark for Measuring Frontier Intelligence of Finance Agents

Aug 2026 · 0 citations · 44 references
Computer Science

TL;DR

Evaluating FrontierFinance, a fully open benchmark of 220 expert-crafted queries and 11,543 source-attributed rubrics spanning six crucial use cases across the full investor workflow, finds that the tool harness, not the model alone, strongly shapes quality and efficiency.

Abstract

AI agents are increasingly deployed for professional investment research, yet no benchmark captures the complexity of the full investor workflow. Existing benchmarks mainly target financial data extraction, a narrow slice that current models have largely saturated, while reference-based metrics and generic LLM-as-a-judge scoring fall short on the open-ended, long-form answers that real analyst queries demand. We introduce FrontierFinance, a fully open benchmark of 220 expert-crafted queries and 11,543 source-attributed rubrics spanning six crucial use cases across the full investor workflow. FrontierFinance is both broader and harder than existing public finance benchmarks. Evaluating frontier models and agent systems under a common harness restricted to publicly available data, we find that the tool harness, not the model alone, strongly shapes quality and efficiency; that Samaya's in-house system leads at 56.0%, ahead of the strongest frontier model (Claude Fable 5, 49.2%) at roughly 2.2x lower cost; and that the best open-weight model (Kimi K3, 46.4%) nearly matches the best proprietary model at 4.5x lower cost. Screening&Discovery and Sector, Industry&Macro remain the hardest use cases across all systems, where even the best systems reach only 33% and 39%. We make the dataset and grading code publicly available.

View source

Similar papers

Jul 2026

Frontier Financial Judgement: Can agents tell what might move a stock?

We introduce Frontier Financial Judgement, a challenging new benchmark developed in collaboration with professional equity analysts to assess agents'ability to replicate expert human judgements. Rapidly identifying new information, evaluating its implications and determining its valuation impact is one of the most time-consuming and challenging aspects of real-world equity coverage. This is becoming ever more difficult and important as AI rapidly increases the quantity of new information to process. The strongest agent we evaluate on Frontier Financial Judgement matches all expert labels in only 52.4% of cases. We also find significant divergence in estimated false-positive rates among frontier agents, ranging from ~1% for GPT-5.6 Sol to ~32% for Claude Sonnet 4.6. To construct the benchmark and make it representative of real-world settings, we combine human-designed and labelled synthetic articles with live news articles and historical documents, creating 656 items for assessment. The resulting task requires agents to distinguish genuinely new, valuation-relevant financial information from stale, immaterial or misleading news under realistic conditions. We find substantial trade-offs among agent accuracy, cost, false positives and reliability that continue to hinder the reliable deployment of news-flow filtering in practice.

Joshua Harris · 0 citations
Review Jul 2026

Capital Markets LLM Reliability Score (CM-LRS): From Plausible to Bankable

This work introduces CM-LRS, a Capital Markets LLM Reliability Score, evaluating outputs at the workflow-output layer across seven dimensions: factual accuracy, evidence traceability, numerical consistency, workflow completeness, source discipline, decision usefulness, and reviewability/auditability.

Prerit Ahuja · 0 citations
Case report Open access Aug 2026

Capable but Not Deployable: Institutional Constraints on AI Exposure in Finance

Empirical measures of AI exposure ask language models to score O*NET tasks for technical feasibility. In finance, technically feasible tasks must still pass through review, documentation, supervision, confidentiality controls, and accountable human sign-off before entering production. We measure the gap between feasibility and institutional deployability using 2,199 O*NET tasks across 99 finance-and-insurance occupations. We score each task with eight frontier models and a prompt ladder that moves from bare capability to finance-industry context and named regulatory regimes. The within-model institutional markdown is about one-fifth of the mean feasibility score, and positive for all eight models. The markdown is largest for regulated, client-facing credit and advice roles and smallest for marketing, software, and support roles. Cross-model agreement also declines as finance context is added: models agree more about what AI can do than about what financial institutions can deploy. Mapping exposure to publicly traded firms through pre-ChatGPT staffing shares, we find that the pricing content resides in the institutional layer: firms in the top half of the markdown distribution underperform the bottom half by roughly 25 percentage points in market-adjusted cumulative abnormal returns over the three years after ChatGPT, while sorting on technical exposure alone produces no gap. The differential lies outside the range the same design produces over every pre-ChatGPT window of equal length, though with one event window and few subsector clusters we read it as evidence on where return information resides rather than as a causal estimate. Especially in regulated industries, deployable exposure rather than technical feasibility is the more relevant measure of AI exposure.

Claes Backman, Christos A. Makridis · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.