Skip to content

FinCUABuild: Can Agents Build Reliable Benchmarks for Dynamic Financial Computer Use?

Sep 2026 · 0 citations · 30 references
Computer Science

TL;DR

This work introduces FinCUABuildBench, a benchmark for evaluating financial CUA task construction, and introduces FinCUABuildAgent, a multi-agent system for automatically constructing dynamic financial CUA evaluation tasks.

Abstract

Financial scenarios are diverse and complex, spanning varying data conditions, tool configurations, and workflows. Yet existing CUA, Computer-Using Agent, evaluation tasks remain largely manually constructed, limiting scalable coverage of real-world financial scenarios. Then, can agents autonomously construct diverse CUA evaluation tasks for financial scenarios? Evaluating this capability poses three key challenges: scenario coverage of construction requests, fair comparison across construction methods, and reliable assessment of generated task quality. To solve these, we introduce FinCUABuildBench, a benchmark for evaluating financial CUA task construction, featuring: (i) 576 construction requests covering 24 financial workflows and three types of runtime variation; (ii) standardized input, budget, and output specifications; and (iii) a task qualification mechanism based on execution tests and quality checks. We further introduce FinCUABuildAgent, a multi-agent system for automatically constructing dynamic financial CUA evaluation tasks. It consists of three modules that jointly construct tasks, environments, and validators. On FinCUABuildBench, under the same model backbone, existing agent-based construction methods achieve strict qualification rates of only 1.3-8.3%, while FinCUABuildAgent reaches 31.3%. Downstream evaluations further show that the constructed tasks can effectively differentiate CUA task-execution capabilities. These results demonstrate that agents can autonomously construct financial CUA tasks with meaningful evaluation value, offering a practical path toward broader evaluation coverage in financial scenarios. Code: https://github.com/FengxianJi/FinCUABuild

View source

Similar papers

Preprint Aug 2026

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks

GDPevo is presented, an evolution-native benchmark grounded in GDP-related enterprise workflows, together with the fully automated data pipeline that generates it, and the best evolved agents remain far below the fully informed oracle ceiling, indicating that the self-evolution ability of current agents remains far fro...

Leijun Zhou, Zhihao Liu, Xiang Qu et al. · 1 citation
Preprint Aug 2026

DuMateBench: Evaluating Autonomous Agents in Complex Real-World Workflows

DuMateBench, a real-session benchmark reconstructed from anonymized and privacy-screened user sessions collected from a large-scale production agent platform, is introduced, showing that performance under environmental perturbations is jointly shaped by the capabilities of the LLM and the surrounding agent framework.

Zechun Niu, Yu-Kun Zhao, Jia-Xin Zhang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

DAREBench: Deployment-Aware and Reliable Evaluation of Models as Agents

As large language models evolve from question-answering systems into general-purpose agents, evaluation must move beyond static answer correctness to assess multimodal perception, multi-step execution, tool use, and artifact delivery. However, existing benchmarks are often tied to specific task types, execution environ...

Yu Liu, Zhi-Lin Liu, Zhi-Wei Yang et al. · 0 citations
#natural language process... Preprint Sep 2026

FinFIRST: Benchmarking Search Agents for Financial Information Retrieval, Sourcing and Traceability

FinFIRST is the first financial benchmark to jointly evaluate answers and supporting evidence through atomic rubrics, retaining final-answer correctness as the primary objective while making the supporting research process measurable, verifiable, and diagnosable.

Wen-Qing Wang, Hai-Tao Xiang, Xin-Yi Zhao et al. · 0 citations
Preprint Aug 2026

AgentMercury: Your Agent Can Synthesize Verifiable Environments for Business Scenarios at scale

This work introduces AgentMercury, a scalable framework for synthesizing executable environments from high-level business scenarios that improves substantially on both enterprise workflows and out-of-domain benchmarks spanning reasoning, coding, scientific computing, and tool use.

Minbyul Jeong, Chanwoong Yoon · 0 citations
Preprint Aug 2026

NetConfArena: An Executable Benchmark for LLM Agents in Closed-Loop Network Configuration

NetConfArena is presented, an executable benchmark for evaluating LLM agents in closed-loop network configuration, and its findings suggest two future directions: using validated trajectories as supervision signals to improve foundation models, and designing harness mechanisms that make agent execution more reliable an...

Chang Liu, Xiao-Hui Xie, Xinyi Chen et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.