Skip to content

CUA-SWE: When Computer-Use Agents Meet Visual Software Engineering

Sep 2026 · 0 citations
Computer Science

TL;DR

This evaluation characterizes how frontier agents combine source-level execution with application screenshots and graphical interaction to produce verified software changes, and examines performance across domains and task information requirements, alongside the development behaviors associated with successful repairs.

Abstract

Software development requires more than editing code: developers repeatedly run software, interact with its interfaces, visually inspect its behavior, and use these observations to decide what to change next and whether a change works. Existing coding agents and computer-use agents are largely studied in isolation, leaving this integrated development process underexplored. Diagnosing a runtime interaction failure requires agents to connect visual observations with the responsible code, then use the application again to verify the repair. We introduce CUA-SWE, a benchmark, environment, and evaluation pipeline for software engineering with computer use. Beyond studying how GUI feedback supports diagnosis and repair, we ask whether agents can complete software engineering tasks when required specification or operational information is available only through the running application's visual interface. CUA-SWE spans four software engineering domains and requires agents to modify code and configuration, execute commands, interact with running software, and inspect visual feedback within the same task. Each task includes deterministic, task-specific tests that verify whether the resulting software satisfies the requirements and preserves specified behavior. Our evaluation characterizes how frontier agents combine source-level execution with application screenshots and graphical interaction to produce verified software changes. We examine performance across domains and task information requirements, alongside the development behaviors associated with successful repairs. CUA-SWE provides a unified testbed for studying how agents use visual feedback and interaction to guide software engineering, with executable correctness criteria for the resulting software.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

WatchPoint: Executable User Feedback for Real-World Agentic Web Development

When a professional web developer's code fails a test, they do not simply re-read the stack trace. They open the application in a browser, click buttons, inspect computed styles, and run diagnostic commands to understand what went wrong. Existing feedback mechanisms for coding agents rely on screenshots, LLM-as-a-judge...

Guanqun Yang, Wei Yang, Xueqing Liu · 0 citations
Review Open access 2026

How Experienced Developers Can Get More Value from Agents

AI agents are becoming a fundamental part of modern software creation, helping developers in generating code, debugging, designing systems, etc. But there is a clear difference between how beginners and experienced software engineers get benefits from these tools. Newbies usually depend on agents for one-time prompts a...

Madhurima Kommuru, Srujana Pulipaka · 0 citations
#artificial intelligence Preprint Sep 2026

The Hard Part Comes After Search: Benchmarking Web Agents on Synthesizing, Organizing, and Displaying Knowledge

Existing computer-use agent benchmarks do not fully evaluate agents acting as assistants. A useful assistant retrieves information across complex, multi-step workflows, synthesizes it into artifacts (documents, presentations, spreadsheets), and navigates program interfaces to produce a coherent final product. Such work...

Alexander Gill, Md Farhan Ishmam, X. Nguyen et al. · 0 citations
#artificial intelligence Preprint Sep 2026

EngiWorld: What Can Frontier Agents Deliver in Professional Engineering Environments?

EngiWorld provides the first rigorous foundation for measuring progress toward agents that operate professional engineering software end to end, and introduces an artifact-centric evaluation methodology built on a unified domain-verifier suite.

Hong-Cheng Gao, Hai-Long Qu, Yu-Ang Lei et al. · 0 citations
Review Aug 2026

Software Engineering for and with GUI Agent

GUI agents have advanced rapidly, producing a growing body of frameworks, benchmarks, and applications. However, this growth has outpaced the maturity of the field. GUI agents remain technically brittle, incompletely engineered, and insufficiently validated for sustained real-world use. They are evolving into closed-lo...

Sheng-Cheng Yu, Yu-Chen Ling, Junyang Xing et al. · 1 citation

RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents

This work introduces RecreationWorld, a five-platform framework built around recreation: given a running reference, an agent must discover its behavior and build a faithful implementation with no prescribed workflow, and introduces RecreationBench, comprising 250 diverse tasks across domains and platforms.

Shuai Bai, Jia-Yong Deng, Si-Cheng Fan et al. · 1 citation

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.