Findings show that functional-only evaluation overestimates agents'ability to satisfy the full requirements of repository-level repair tasks, and introduces SWE-Gate, a repository-level benchmark for software engineering agents that explicitly evaluates review constraint compliance alongside functional correctness.
Xin He, Yan-Lin Wang, Ming-Wei Liu et al.· 0 citations
PTA-IRT is proposed, a Privileged Trajectory-Aware Item Response Theory framework that fuses process and outcome signals and consistently outperforms prior IRT baselines on score and ranking recovery across four SWE benchmarks.
Ke-Feng Duan, De-Wu Zheng, Yan-Lin Wang et al.· 0 citations
PhoenixRepair is a multi-agent framework that systematically explores multiple candidate edit locations and performs iterative reflection and refinement on patch generation, thereby expanding the search space of repair strategies and achieves higher fault localization accuracy than existing approaches.
Tian-Yue Jiang, Yan-Lin Wang, Xin He et al.· arXiv.org· 2 citations
This work proposes WebDesignIter, a framework built around a persistent knowledge graph (WebAppArchKG) that fuses repository structure with design knowledge and keeps both in sync across development cycles, and outperforms every general-purpose coding agent Claude Code, OpenHands, SWE-Agent, Codex CLI on every model co...
RepoReasoner is introduced, a benchmark for evaluating repository-level code reasoning that assesses two complementary abilities: Output Prediction, which measures fine-grained, stateful execution reasoning across files, and Call Chain Prediction, which evaluates high-level architectural dependency understanding under...
Yanlin Wang, Suiquan Wang, Yanlin Wang et al.· Proceedings of the ACM on So...· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.