E4R-Reviewer: Effective and Explainable Automated Code Review via End-to-End Reasoning-Guided Alignment
Abstract
Code review is a key practice for ensuring software quality and maintainability. Despite progress in Automated Code Review (ACR), existing methods face two core challenges: (1) Isolated Task Modeling. Current approaches often model and optimize subtasks in ACR independently, ignoring the inherent logical order and internal dependencies among them which affects the effectiveness of ACR. (2) Lack of Explainability. At the task level, the absence of explanatory information in review comments increases developers’ cognitive load; at the model level, the black-box nature fundamentally undermines developer trust. To address these challenges, we propose E4R-Reviewer, which improves the Effectiveness and Explainability of ACR through End-to-End Reasoning-guided alignment. For effectiveness, E4R-Reviewer unifies multiple fine-grained ACR subtasks into a single end-to-end reasoning process, enabling cross-task knowledge sharing and allowing the model to explicitly complete a reasoning chain that covers quality estimation, issue localization, issue classification, issue description, fix suggestion, and code refinement in one generation. Meanwhile, we adopt a Group Relative Policy Optimization (GRPO)-based reinforcement-learning alignment, treating the reasoning steps as optimizable intermediate objectives. We design subtask-specific rewards and integrate them via curriculum-inspired, multi-stage reward fusion that follows the real-world review workflow. For explainability, E4R-Reviewer produces reasoning process and structured review results covering all fine-grained ACR subtasks, improving the transparency and explainability of the review results. Extensive evaluations on public, real-world datasets demonstrate that E4R-Reviewer significantly outperforms existing methods and achieves state-of-the-art performance: a 74.61% F1-score in quality estimation and +22.96% CodeBLEU in code refinement. Furthermore, Large Language Model (LLM) and human evaluation further confirm the superiority of E4R-Reviewer in terms of effectiveness and explainability.