Skip to content

Author

Yanlin Wang

We have 5 of 63 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Review Sep 2026

SWE-Gate: Passing Functional Tests Is Not Enough for Software Engineering Agents

Findings show that functional-only evaluation overestimates agents'ability to satisfy the full requirements of repository-level repair tasks, and introduces SWE-Gate, a repository-level benchmark for software engineering agents that explicitly evaluates review constraint compliance alongside functional correctness.

Xin He, Yan-Lin Wang, Ming-Wei Liu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Efficient SWE Agent Benchmarking via Trajectory-Aware Evaluation

PTA-IRT is proposed, a Privileged Trajectory-Aware Item Response Theory framework that fuses process and outcome signals and consistently outperforms prior IRT baselines on score and ranking recovery across four SWE benchmarks.

Ke-Feng Duan, De-Wu Zheng, Yan-Lin Wang et al. · 0 citations
Jul 2026

PhoenixRepair: Rethinking Repair Strategy Exploration in Software Agents

PhoenixRepair is a multi-agent framework that systematically explores multiple candidate edit locations and performs iterative reflection and refinement on patch generation, thereby expanding the search space of repair strategies and achieves higher fault localization accuracy than existing approaches.

Tian-Yue Jiang, Yan-Lin Wang, Xin He et al. · 2 citations
Review Jul 2026

WebDesignIter: Co-Evolving Design Knowledge for Repository-Level Front-End Code Generation

This work proposes WebDesignIter, a framework built around a persistent knowledge graph (WebAppArchKG) that fuses repository structure with design knowledge and keeps both in sync across development cycles, and outperforms every general-purpose coding agent Claude Code, OpenHands, SWE-Agent, Codex CLI on every model co...

Zheng Pei, Mingwei Liu, Zhenxi Chen et al. · 1 citation · ⚡1
Open access Jun 2026

RepoReasoner: Evaluating Repository-Level Code Reasoning Ability of Long-Context Language Models

RepoReasoner is introduced, a benchmark for evaluating repository-level code reasoning that assesses two complementary abilities: Output Prediction, which measures fine-grained, stateful execution reasoning across files, and Call Chain Prediction, which evaluates high-level architectural dependency understanding under...

Yanlin Wang, Suiquan Wang, Yanlin Wang et al. · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.