This evaluation characterizes how frontier agents combine source-level execution with application screenshots and graphical interaction to produce verified software changes, and examines performance across domains and task information requirements, alongside the development behaviors associated with successful repairs.
Abstract
Software development requires more than editing code: developers repeatedly run software, interact with its interfaces, visually inspect its behavior, and use these observations to decide what to change next and whether a change works. Existing coding agents and computer-use agents are largely studied in isolation, leaving this integrated development process underexplored. Diagnosing a runtime interaction failure requires agents to connect visual observations with the responsible code, then use the application again to verify the repair. We introduce CUA-SWE, a benchmark, environment, and evaluation pipeline for software engineering with computer use. Beyond studying how GUI feedback supports diagnosis and repair, we ask whether agents can complete software engineering tasks when required specification or operational information is available only through the running application's visual interface. CUA-SWE spans four software engineering domains and requires agents to modify code and configuration, execute commands, interact with running software, and inspect visual feedback within the same task. Each task includes deterministic, task-specific tests that verify whether the resulting software satisfies the requirements and preserves specified behavior. Our evaluation characterizes how frontier agents combine source-level execution with application screenshots and graphical interaction to produce verified software changes. We examine performance across domains and task information requirements, alongside the development behaviors associated with successful repairs. CUA-SWE provides a unified testbed for studying how agents use visual feedback and interaction to guide software engineering, with executable correctness criteria for the resulting software.
When a professional web developer's code fails a test, they do not simply re-read the stack trace. They open the application in a browser, click buttons, inspect computed styles, and run diagnostic commands to understand what went wrong. Existing feedback mechanisms for coding agents rely on screenshots, LLM-as-a-judge...
AI agents are becoming a fundamental part of modern software creation, helping developers in generating code, debugging, designing systems, etc. But there is a clear difference between how beginners and experienced software engineers get benefits from these tools. Newbies usually depend on agents for one-time prompts a...
Madhurima Kommuru, Srujana Pulipaka· International Journal of Mod...· 0 citations
Existing computer-use agent benchmarks do not fully evaluate agents acting as assistants. A useful assistant retrieves information across complex, multi-step workflows, synthesizes it into artifacts (documents, presentations, spreadsheets), and navigates program interfaces to produce a coherent final product. Such work...
Alexander Gill, Md Farhan Ishmam, X. Nguyen et al.· 0 citations
EngiWorld provides the first rigorous foundation for measuring progress toward agents that operate professional engineering software end to end, and introduces an artifact-centric evaluation methodology built on a unified domain-verifier suite.
Hong-Cheng Gao, Hai-Long Qu, Yu-Ang Lei et al.· 0 citations
GUI agents have advanced rapidly, producing a growing body of frameworks, benchmarks, and applications. However, this growth has outpaced the maturity of the field. GUI agents remain technically brittle, incompletely engineered, and insufficiently validated for sustained real-world use. They are evolving into closed-lo...
Sheng-Cheng Yu, Yu-Chen Ling, Junyang Xing et al.· 1 citation
This work introduces RecreationWorld, a five-platform framework built around recreation: given a running reference, an agent must discover its behavior and build a faithful implementation with no prescribed workflow, and introduces RecreationBench, comprising 250 diverse tasks across domains and platforms.
Shuai Bai, Jia-Yong Deng, Si-Cheng Fan et al.· 1 citation
With $2.1 million funding from Google.org, the open-source Public Transit Intelligence Hub will unify public transit monitoring, operations, and passenger communication.
Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.