FinDeepIndicator is proposed, the first benchmark dedicated to evaluating Deep Research agents in end-to-end financial indicator construction, and covers fundamental, technical, and macroeconomic indicators organized into 21 fine-grained sub-categories.
Abstract
Financial indicators are essential tools for transforming raw financial data into interpretable measures for various downstream tasks, such as valuation, risk assessment, and economic analysis. However, existing financial benchmarks largely focus on answer-level accuracy and often assume that relevant data are already provided, leaving the assessment of the intermediate process of indicator construction underexplored. In this work, we propose FinDeepIndicator, the first benchmark dedicated to evaluating Deep Research (DR) agents in end-to-end financial indicator construction. Specifically, FinDeepIndicator evaluates DR agents across four stages in indicator construction: formula specification, data collection, indicator calculation, and answer generation, and covers fundamental, technical, and macroeconomic indicators organized into 21 fine-grained sub-categories. It contains 3,350 curated question-answer (QA) pairs derived from both U.S. and Chinese markets, 10 years of historical financial data, and 800 listed companies. Extensive experiments on search-equipped Large Language Models (LLMs) and DR agents show that, while LLMs generally perform well in formula specification, their accuracy drops substantially during data retrieval and numerical execution. DR agents consistently outperform search-equipped LLMs, yet remain unreliable in realistic financial analysis settings. These findings provide insights for developing more capable and trustworthy DR agents in finance.
QFRS underpins a public state-of-the-art leaderboard, ensuring that only studies satisfying these standards are ranked, with the goal of shifting the literature from opaque, error-metric-driven results to transparent, economically meaningful and comparable benchmarks.
M. Khushi· Artificial Intelligence Revi...· 1 citation
Evaluating FrontierFinance, a fully open benchmark of 220 expert-crafted queries and 11,543 source-attributed rubrics spanning six crucial use cases across the full investor workflow, finds that the tool harness, not the model alone, strongly shapes quality and efficiency.
Yuhao Zhang, O. O. Koyluoglu, Thejas Venkatesh et al.· 0 citations
Powered by advances in LLMs and autonomous agents, deep research has become one of the most widely adopted agentic products. However, most deep research systems write general-purpose reports, which are inadequate for financial deep research. Financial research demands specialized knowledge to analyze historical patterns and forecast upcoming events. Automating financial deep research therefore requires both a layered harness to drive the research agent and a verifiable, point-in-time benchmark that prevents leakage of future information. We present FinanceHarness, a harness that runs finance-oriented tools and practitioner-guided workflows, automating financial deep research end to end: environment and data construction, the agent execution loop, and reward modeling. We further propose FinanceGym, comprising thesis-driven research questions and rubrics that combine pre-cutoff and post-cutoff criteria. Professional expert validation yields an 82% pass rate. With the same open-weight backbone, FinanceHarness improves the overall rubric score from 25.3% to 32.4%, demonstrating the effectiveness of our specialized harness design. However, even pairing FinanceHarness with the most cutting edge LLM (e.g. Opus-5), the FinanceGym score is below 45%, showing that it is a challenging benchmark for financial deep research. Leaderboard is available at: https://financegym.github.io/ and FinanceHarness code is available at: https://github.com/Yijia-Xiao/FinanceHarness.
Research on stock market prediction relies on public datasets to develop and compare learning-based models. However, existing datasets are often limited to price-based information, cover a restricted set of assets or time periods, or provide only a subset of the heterogeneous signals required by modern prediction approaches. Moreover, differences in data preprocessing, task formulation, and evaluation protocols make experimental results difficult to reproduce and hard to compare across studies. To address these challenges, we present FinBench, a benchmarking framework designed to support reproducible evaluation of stock market prediction models using heterogeneous financial data. FinBench integrates historical prices, inter-company relations, macroeconomic indicators, and financial news within a unified data construction pipeline, and standardizes feature extraction and evaluation across classification, regression, and ranking tasks. Model outputs are evaluated both at the task prediction level and through portfolio-based strategies under consistent experimental settings. Using FinBench, we benchmark a range of existing models on European and US equity markets, analyzing their behavior across different prediction tasks, portfolio constructions, and investment universes with diverse liquidity and structural characteristics. FinBench provides a practical reference for systematic and reproducible comparison of stock market prediction methods.
Sara Pederzoli, Marta Santacroce, Francesco Guerra et al.· Proceedings of the 32nd ACM...· 0 citations
AnalysisBank is proposed, which distills expert reports into a reusable library of Analyses, each pairing a data signal, an analytical move, and the expert span it was derived from, to operate at the analytical rather than structural level of financial report generation.
Ya-Jing Yang, Yunshan Ma, Kelvin J. L. Koa et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.