Skip to content

SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents

Sep 2026 · 1 citation · 38 references
Computer Science

TL;DR

SWE-Bench Pro Verified offers a more trustworthy benchmark for assessing software engineering agents, which combines anti-hacking safeguards that eliminate major leakage channels without disrupting normal agent functionality, with task refinement that minimally corrects inconsistencies within flawed instances.

Abstract

SWE-Bench Pro has emerged as a standard benchmark for evaluating software engineering agents on challenging repository-level tasks. However, our analysis work show that its evaluation is undermined by two sources of unreliability: reward hacking, enabled by leakage of gold solutions or hidden evaluation information, and task quality issues, including misleading problem statements and improperly scoped tests. These issues can inflate benchmark performance and obscure agents'true coding ability. We present SWE-Bench Pro Verified, a verified version of SWE-Bench Pro that addresses both problems. Our approach combines anti-hacking safeguards that eliminate major leakage channels without disrupting normal agent functionality, with task refinement that minimally corrects inconsistencies within flawed instances. Evaluations on SWE-Bench Pro Verified reveal that some models perform substantially worse than previously evaluated, suggesting that existing results on SWE-Bench Pro may overestimate real software engineering capability. SWE-Bench Pro Verified offers a more trustworthy benchmark for assessing software engineering agents.

View source

Similar papers

#artificial intelligence Review Sep 2026

SWE-Gate: Passing Functional Tests Is Not Enough for Software Engineering Agents

Findings show that functional-only evaluation overestimates agents'ability to satisfy the full requirements of repository-level repair tasks, and introduces SWE-Gate, a repository-level benchmark for software engineering agents that explicitly evaluates review constraint compliance alongside functional correctness.

Xin He, Yan-Lin Wang, Ming-Wei Liu et al. · 0 citations
Review Aug 2026

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring

SWE-Bench ProMax is introduced, an expert-curated, multilingual code refactoring benchmark of 170 instances drawn from real commits across seven programming languages, which presents a meaningful and unsaturated challenge for current AI coding agents.

Yu-Ling Shi, Jing-Heng Xu, Kelin Fu et al. · 6 citations
#software testing Preprint Aug 2026

SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?

SWE Refactor Bench is introduced, a benchmark comprising 20 whole-repository migrations, covering 4 kinds of technical debt, and SWE Refactor Bench is positioned as a rigorous testbed for developing coding agents for reliable whole-repository migrations.

De-Yao Hong, Yi-Zhe Chi, Wen-Yi Li et al. · 3 citations · ⚡1
#software testing Preprint Sep 2026

A Large-Scale Empirical Study of Quality Assurance Practices and Gaps in AI Agents

A large-scale empirical study of quality assurance (QA) practices in 157 open-source LLM-based agent projects with at least 100 GitHub stars highlights the need to move beyond feature-level testing toward systematic end-to-end validation that ensures agent workflows remain within intended boundaries when interacting wi...

Wu-Yang Dai, Moses Openja, Jiho Shin et al. · 0 citations
Preprint Aug 2026

SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?

A paired ablation that removes explicit scientific guidance while preserving the repository and executable engineering context shows that scientific knowledge is not uniformly beneficial: well-grounded information can constrain repair and improve average performance and token efficiency, whereas poorly aligned guidance...

Zhi-Peng Xu, Jia-Hao Lu, Yi-Ning Zheng et al. · 4 citations
#artificial intelligence Preprint Sep 2026

SWE-Prometheus: Measuring Engineering Governance Improvements in Real-World Repositories

Large language model based coding agents have made substantial progress on repository-level software engineering tasks. Existing repository benchmarks, however, usually start from a human-identified issue and evaluate whether a patch satisfies a functional signal. We present SWE-Prometheus, a benchmark for the broader...

Jia-Jun Wu, Lei-Xin Sun, Zi-Hang Tan et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Oct 2, 2026

Documenting the tech worker movement

Writing as a participant and researcher, PhD student JS Tan SM ’22 has co-authored a new book about the rise of tech worker protests and the employer backlash that followed.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.