Skip to content
Preprint

CorporateBench: Large-Scale Q&A Benchmarking with Temporal Knowledge Bases

Aug 2026 · 0 citations · 57 references
Computer Science

TL;DR

CorporateBench (CB), a human-validated multi-task Q&A benchmark whose scale approaches the conditions LLMs encounter in corporate communication networks, is presented, with evaluation corpora surpassing 230,000 documents.

Abstract

LLMs are increasingly able to answer complex questions about enterprise-scale document collections. But evaluation is hard: companies don't want to share internal communications, and synthetic datasets have been overly simple. We present CorporateBench (CB), a human-validated multi-task Q&A benchmark whose scale approaches the conditions LLMs encounter in corporate communication networks, with evaluation corpora surpassing 230,000 documents. CB evaluates LLMs across two dimensions (information extraction and knowledge base querying) through four synthetically generated firms ranging from 12 to 10,000 employees. Each corpus is sampled from a temporally evolving knowledge base describing a consistent world, guaranteeing cross-document logical consistency even across hundreds of thousands of documents. We evaluate five LLMs on CB, revealing increasingly poor performance as input size approaches realistic scales. CB provides LLM developers a metric for corporate communication reasoning, filling a crucial gap in the benchmarking ecosystem.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

DI-Bench: Systematically Generating In-Domain Data Intelligence Benchmarks for Enterprise Agents

Evaluating enterprise agents on domain-specific benchmarks is critical, yet public benchmarks rarely evaluate whether agents can integrate business knowledge with analytical computation, and constructing such benchmarks manually is costly. We present DI-Bench, a pipeline for generating realistic benchmarks for data int...

Jiang-Yun Zhang, K. Surrao, Torpong Nitayanont et al. · 0 citations
#natural language process... Preprint Sep 2026

FinFIRST: Benchmarking Search Agents for Financial Information Retrieval, Sourcing and Traceability

Financial search is a highly demanding task for LLM agents, requiring not only a correct final answer but also temporally valid information retrieval, authoritative source selection, entity and period alignment, unit and definition consistency, and verifiable evidence for all conclusions. Existing benchmarks predominan...

Wen-Qing Wang, Hai-Tao Xiang, Xin-Yi Zhao et al. · 0 citations
Preprint Aug 2026

ENTLORE: A Graph-Grounded Benchmark for Latent Organizational Reasoning in Enterprise Question Answering

Enterprise question answering is framed as retrieving internal documents and generating grounded answers. Routine enterprise records, however, are work by-products in which required organizational relations remain implicit across heterogeneous sources. Existing benchmarks provide realistic multi-source evidence, but of...

Akrin Zheng, Alexander Wu, Alaia Liu · 1 citation · ⚡1
#artificial intelligence Preprint Aug 2026

FinRCA-Bench: Benchmarking Evidence Retrieval and Reasoning for Financial AI Systems

Large language models are increasingly used to support financial operations, but their apparent reasoning performance can depend on whether they receive the right evidence. In financial reconciliation, the evidence needed for diagnosis is distributed across invoices, purchase orders, approvals, allocations, payments, l...

Prof. S. B. Ghawate · 1 citation
Review Aug 2026

Benchmarking and Enhancing LLMs for Rule-Intensive Review of National Standard Documents

Large language models (LLMs) increasingly support complex professional tasks, yet their capabilities in rule-intensive document review remain insufficiently evaluated. National standard documents, such as China GB/T standards, offer a representative testbed: they are lengthy, highly structured, and governed by explicit...

Tao Wang, Qihao Yang, Rongjiao Liang et al. · 0 citations
Book Open access Aug 2026

CEComBench: Benchmarking Large Language Models' performance on Chinese E-commerce tasks

We introduce CEComBench (Chinese E-Commerce Benchmark), a rigorously curated evaluation framework comprising 12140 annotated samples spanning 36 distinct tasks, sourced from JD.com, a leading Chinese e-commerce platform. Crucially, our data collection, task generation, and evaluation pipeline eschew LLM involvement to...

Guangtao Nie, Huimu Wang, Gewei Lu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.