Skip to content
Preprint

FinRank: An Evidence-Grounded Benchmark for Financial Question Answering and Retrieval over SEC Filings

Aug 2026 · 0 citations · 35 references
Computer Science Economics

TL;DR

FinRank provides an evidence-first benchmark for developing financial question answering systems that are not only accurate but also grounded in the correct disclosure.

Abstract

Financial question answering is typically evaluated by answer correctness, yet in SEC filings a plausible and even numerically correct answer can be grounded in the wrong evidence. Similar facts and disclosures recur across sections of a filing, across reporting periods of the same firm, and across comparable firms. FinRank targets this provenance-sensitive retrieval problem by requiring systems to identify evidence for the intended entity, reporting period, and disclosure context. The benchmark contains 1185 manually authored question-answer records over the 10-K and 10-Q filings of 22 companies. Each record includes a reference answer, gold supporting passages, and hand-curated hard negatives drawn from confusable passages within filings, across reporting periods, and across comparable firms. FinRank evaluates passage retrieval, reranking, and hard-negative discrimination as separately measured tasks. Baseline results demonstrate the difficulty of this setting: among the evaluated systems, even a 7B instruction-tuned embedder reaches only 44.8% Recall@10 on the pooled evidence corpus; sub-billion-parameter encoders gain at most 3.5 points over BM25, a finance-adapted embedder trails BM25 by 9.7 points, and pairwise accuracy falls by 13.0-20.5 percentage points when random negatives are replaced with the curated hard negatives. FinRank provides an evidence-first benchmark for developing financial question answering systems that are not only accurate but also grounded in the correct disclosure.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Enhancing Financial Question Answering: A Novel Benchmark Dataset of Banks'financial statements

FinRAG-QA is introduced, a novel benchmark dataset for financial question answering, which comprises 999 practitioner-curated questions on 10 standardised indicators, grounded in 209 annual and Pillar 3 reports from 24 major European and U.S. banks spanning 2019-2023.

Arianna Miola, Bruno Spaccavento, Lorenzo Silotto et al. · 0 citations
Preprint Aug 2026

HNR-DAC: Hard-Negative Reranking and Distribution-Aligned Classification for Scientific Claim Verification

This work presents HNR-DAC, a two-stage framework that trains each stage on the cases it will actually encounter, and quantifies evidence confusability using a base reranker's scores on non-gold paragraphs and contrasts gold evidence against the most confusable candidates.

Zhenchao Wang, Xin Chen, Luo-Xiu Zhang et al. · 0 citations
Jul 2026

FinSAgent: Corpus-Aligned Multi-Agent RAG Framework for Evidence-Grounded SEC Filing Question Answering

Financial question answering over U.S. Securities and Exchange Commission (SEC) filings requires retrieving and synthesizing heterogeneous evidence dispersed across long, standardized, and highly redundant disclosures. Existing retrieval-augmented and multi-agent systems typically derive retrieval queries directly from...

Jijun Chi, Zhenghan Tai, Hanwei Wu et al. · 0 citations
#small language model Open access Aug 2026

Retrieval Granularity as Evidence Design in Small-Model RAG Question Answering: A Diagnostic HotpotQA Study

Results align with a diagnostic perspective on chunking: using evidence at a task-appropriate level of granularity can improve grounding, auditability, and answer quality, but the observed patterns should be interpreted within the HotpotQA distractor setting, fixed generator, and tested context budgets.

Wei-Mao Ke, Li-Xia Yang, Meng-Yang Xu · 0 citations
Review Jul 2026

Overview of FinMMEval 2026 Task 2: Multilingual Financial Short-Answer Question Answering

FinMMEval 2026 Task 2 evaluates short-answer financial question answering over multilingual evidence over multilingual evidence using document retrieval-augmented generation, cross-lingual evidence handling, structured prompting, answer compression, and validation strategies.

Zhuohan Xie, Xueqing Peng, Georgi N. Georgiev et al. · 1 citation
Preprint Aug 2026

When Absence Is Evidence: Evaluating Completeness-Sensitive Negative Reasoning in Large Language Models

CROWN-QA is introduced, comprising CROWN-Synth, a controlled paired core that fixes the question and observed facts while varying only query-relative coverage, and CROWN-Real, a real-document contrast-set evaluation with controlled coverage variants.

Byoungjae Min, Kennedy Edemacu, Sae-Hong Cho et al. · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.