#machine learning
Jun 2026
Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-style Agent Harnesses on Coding Tasks
Claw-SWE-Bench is introduced, a unified benchmark that enables researchers to systematically assess the capabilities and efficiency of general-purpose harnesses on software engineering tasks and guide the development of more capable and efficient harnesses.
Meng-Yu Zheng, Kai Han, Bo-Xun Li et al.
· arXiv.org · 5 citations