This evaluation characterizes how frontier agents combine source-level execution with application screenshots and graphical interaction to produce verified software changes, and examines performance across domains and task information requirements, alongside the development behaviors associated with successful repairs.
P. Wang, Chen-Hao Liang, Ze-Long Xu et al.· 0 citations
WeClawArena is introduced, an auditable benchmark and runtime sandbox for multi-party owned-agent collaboration over personal workspaces and audits attack success from bounded runtime evidence, supporting diagnosis of task breakdown, privacy leakage, poisoned evidence, and invalid authority paths.
P. Wang, Ao-Jie Yuan, Haiyu Zhang et al.· 1 citation
CatchBench puts one auditor's question to three information states: the declared configuration before a run (PRE), a growing prefix of its trace (LIVE), and the finished trace (POST), which none scores all three under one task-method interface.
Yue Zhao, Meng-Yuan Li, Ruo-Lin Li et al.· 4 citations· ⚡1
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.