Skip to content
Review

SWE-Review: Closing the Loop on Issue Resolution with Agentic Code Review

Jul 2026 · arXiv.org · Vol abs/2607.06065 · 4 citations · 47 references
Computer Science

TL;DR

Experiments show that agentic review continuously improves PRs through a generate-review-revise loop, outperforms single-turn fixed-context review in both decision accuracy and resolve rate after revision, transfers beyond review to improve issue-resolution models, and enables effective and efficient test-time scaling.

Abstract

Coding agents increasingly generate pull requests (PRs) for real-world software issues, yet one-shot PR generation remains open-loop: the PR is proposed without systematic review, diagnosis, or revision. We introduce \textbf{SWE-Review}, a framework for closing this loop with agentic code review. Given an issue and an AI-generated PR, a reviewer agent explores the repository, decides whether the PR should be accepted, and provides structured feedback for revision. We evaluate this setting with our proposed \textbf{SWE-Review-Bench} to measure both review correctness and downstream revision usefulness. We further curate \textbf{SWE-Review-Traj} dataset to study broader applications of agentic review and fill the data-scarcity gap for open reviewer training. Experiments show that agentic review continuously improves PRs through a generate-review-revise loop, outperforms single-turn fixed-context review in both decision accuracy and resolve rate after revision, transfers beyond review to improve issue-resolution models, and enables effective and efficient test-time scaling. These results position agentic code review as a practical mechanism for moving AI coding agents from one-shot PR generation toward closed-loop issue resolution.

View source

Similar papers

Review Jul 2026

"Go Home Copilot, You're Drunk": Understanding Developer Responses to Agent-Generated Code Review Comments

The first large-scale empirical study on the resolution of agent-generated code review comments is presented, revealing that the presence of an inline code suggestion is the strongest predictor of comment resolution, while lengthy and complex comments are less likely to be acted upon.

Shamse Tasnim Cynthia, Ratnadira Widyasari, Banani Roy et al. · 0 citations
Review Aug 2026

OpenCodeReview: Determinism over Non-Determinism for Cost-Effective Agent-Based Code Review

OpenCodeReview is introduced, built on deterministic engineering for uncertain agents: rather than granting maximal freedom, determinism is injected at three deliberate pipeline points to address non-determinism in LLM code review systems.

Zhengfeng Li, Lei Zhang, Xian Wu et al. · 0 citations
Review Jul 2026

Cross-Model LLM Code Review: Should you use Claude to review Codex or vice versa?

Developers increasingly use two coding agents together: one writes a draft, and the other reviews it. However, it is not clear whether the pairing is worth its cost and time, or whether the order of the pairing matters. We run a controlled experiment on 116 recent hard and medium lcb tasks with Claude and Codex across...

Zuodong Xiang, Yi-Ke Zhang, Yue-Ming Zhang et al. · 1 citation
Jul 2026

Reflection or Re-Generation? Why LLM Revision Fails Where Human Revision Succeeds

The Human-LLM Reflection Framework is introduced, a controlled two-pass protocol comparing human and LLM revision under identical conditions across self-, peer-, and cross-agent settings, using an information-theoretic analysis based on per-iteration cross-entropy reduction.

Ye-Fan Tao, Gerald Friedland, Madhusudhanan Chandrasekaran et al. · 0 citations
Review Jul 2026

Code Review is a Conversation: Toward Conversational AI Review Assistants

This vision reframes AI code review from automated commenting to human-AI sensemaking before integration, and outlines a research agenda for studying review conversations, designing conversational AI review capabilities, and evaluating their impact on software evolution and maintenance.

Rosalia Tufano · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.