Skip to content

ERPBench: A State-Grounded Evaluation Paradigm for Computer-Use Agents in Enterprise Software

Sep 2026 · 1 citation · 18 references
Computer Science

TL;DR

ERPBench is introduced, a benchmark that evaluates screenshot-only agents on a live and reproducible system and scores each task against ground-truth values in its database and presents a production-grade harness that gates agent actions behind human approval for safe deployment.

Abstract

Computer-use agents that operate through screenshots and simulated actions are advancing rapidly, yet their evaluation remains anchored to general desktop and web tasks. Enterprise Resource Planning systems run the finance, procurement, inventory, and customer operations of organizations worldwide, and pose distinct challenges for computer-use agents: dense interfaces, coordinated multi-step interactions, and errors that alter persistent business records rather than surfacing on screen. Existing enterprise computer-use benchmarks rely on proprietary platforms or on simulated approximations of such software. We introduce ERPBench, a benchmark that evaluates screenshot-only agents on a live and reproducible system and scores each task against ground-truth values in its database. Beyond the benchmark, we present a production-grade harness that gates agent actions behind human approval for safe deployment. Evaluating six closed and open-source agents, we demonstrate that strong general performance does not transfer to enterprise reliability. Even when an agent reaches the right form and saves it, the stored record is often wrong: some agents save in up to 85% of runs but write the correct value in as few as 3%. We further characterize failure modes specific to enterprise workflows.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

$\tau^\tau$-Bench: An Environment for End-To-End, Realistic Agent Construction

LLM agents are rapidly becoming production software, deployed to handle customer service, adjudicate disputes, and operate internal systems. Notably, the work of building them is increasingly handed to coding agents, yet existing benchmarks say little about whether an AI system can deliver one under the conditions of a...

Shi-Xiu Quan, Keshav Dhandhania, Karthik R. Narasimhan et al. · 1 citation
Review Aug 2026

Evaluating Skills, Not Just Agents: Agentic Continuous Evaluation of Skills

ACES (Agentic Continuous Evaluation of Skills), a repository-native framework for evaluating skills and product capability packages as executable agent artifacts, is presented.

Christopher Kevin, Narendran Raghavan, J. Puget et al. · 3 citations · ⚡1
#artificial intelligence Preprint Sep 2026

The Agentic Company OS: Substrate Inversion for Sustained Enterprise Agent Deployment

Enterprise AI agents often succeed in a demonstration and then stall once they must operate day after day. An industry report estimates that most pilots never reach production and that deployed systems rarely retain feedback or improve over time, while agent benchmarks show single-run successes masking unreliable repet...

Oliver Aleksander Larsen, M. T. Moghaddam · 0 citations
#artificial intelligence Preprint Aug 2026

ComponentBench: Diagnosing Component-Level Failures in Computer-Use Agents

This work presents ComponentBench, a benchmark and diagnostic pipeline for component-level evaluation of computer-use agents on modern web UIs, and introduces a scalable pipeline for auditing realized structural difficulty after implementation and synthesizing structured failure analyses across tasks and component fami...

Tian-Chen Guan, Xinlei Lin, Royce Cheng-Yue et al. · 0 citations
Review Aug 2026

Terminal Agents: A Survey of AI Agents in Command-Line Environments

This survey establishes workload-level boundaries and connects system architecture, competence acquisition, and evaluation through a seven-dimensional terminal competence profile, and provides a unified basis for studying terminal-mediated agency across software engineering and emerging application domains.

Yi Bin, Xiao-Yang Yuan, Hao Zeng et al. · 1 citation

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.