An empirical study of agentic code reviews using CodeRabbit as a case study shows that agentic reviews receive mixed reception, and finds that lightweight learning-based methods achieve up to 76% F1 score, suggesting learnable patterns exist between code reviews and their corresponding feedback.
Abstract
Agentic code review, where autonomous agents provide code review comments on pull requests, is increasingly integrated into development workflows, yet there is limited empirical evidence on how developers respond to such comments in practice. In this paper, we present an empirical study of agentic code reviews using CodeRabbit as a case study. Through an empirical study of 31,073 pairs of code reviews and developer feedback from 10,191 pull requests across 239 GitHub repositories, our results show that agentic reviews receive mixed reception: 36.4% were accepted and 7.3% triggered discussion, while 56.3% were rejected. Rejections were primarily associated with invalid suggestions that were false positives, redundant, or out of scope, as well as misalignment with developer intent and coding practices. We further found that agentic reviews tend to focus more on functional concerns than evolvability-related comments, yet they were more likely to be invalid. To improve effectiveness in review practices, we explored various LLM-based approaches for predicting review rejection. We found that lightweight learning-based methods achieve up to 76% F1 score, suggesting learnable patterns exist between code reviews and their corresponding feedback. Our results highlight the current state of CodeRabbit's agentic code reviews, showing opportunity gaps for improvement, as well as shortcomings hindering its effectiveness.
The first large-scale empirical study on the resolution of agent-generated code review comments is presented, revealing that the presence of an inline code suggestion is the strongest predictor of comment resolution, while lengthy and complex comments are less likely to be acted upon.
Shamse Tasnim Cynthia, Ratnadira Widyasari, Banani Roy et al.· arXiv.org· 0 citations
The results show that agent-involved collaboration patterns, especially reviews initiated by AI agents or involving multiple AI agents, are associated with faster review decisions under Gradual AI Adoption and Rapid AI Agent Adoption, but these efficiency gains do not translate into better review quality.
Suzhen Zhong, Shayan Noei, Bram Adams et al.· 2 citations
This study investigates how large language models can assess code review feedback quality along two dimensions, sentiment and specificity, to support more constructive collaboration, and demonstrates the feasibility and practical utility of automated feedback-quality assessment in real-world environments.
Audrey You, J. Wang, You-Xiang Lei et al.· SIGSOFT FSE Companion· 0 citations
It is found that a good bug report for an agent overlaps with, but is not identical to, a good report for a human: agents benefit most from concrete, executable, and well-localized information, whereas some qualities long emphasized for human readers, such as natural language steps to reproduce and readable description...
Lara Khatib, N. Mathews, M. Nagappan et al.· arXiv.org· 2 citations· ⚡1
This vision reframes AI code review from automated commenting to human-AI sensemaking before integration, and outlines a research agenda for studying review conversations, designing conversational AI review capabilities, and evaluating their impact on software evolution and maintenance.
A high-quality benchmark of 1,000 code refinement instances from 328 Python, Java, and JavaScript repositories that focused on one of the most challenging code refinement scenarios that strictly requires repository-level knowledge reasoning, and a straightforward method, RepoRefiner, which retrieves repository-level co...
Ke Wang, Peng Lan, Jia-Kun Liu et al.· ACM Transactions on Software...· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.