It is argued that improving review comment generation requires more than dataset cleaning alone, motivating explicit validity criteria, richer contextual inputs, and evaluation practices aligned with review intent and actionability.
Abstract
Generating code review comments has become a prominent research direction in automated code review, commonly formulated as a text generation task over diff-comment pairs. Despite advances in learning-based approaches, generated review comments are often generic, weakly grounded, or non-actionable. Recent studies have also shown that review comment datasets contain noisy or unsuitable training instances, motivating LLM-based dataset cleaning approaches. In this paper, we argue that problematic training instances are not homogeneous and that some limitations stem from deeper issues in the task formulation itself. Through an empirical inspection of a widely used review comment dataset, we identify misaligned training pairs: instances where the relationship between the code change and the review comment does not provide a reliable learning signal for generating actionable review feedback from localized inputs. We derive a taxonomy of misalignment capturing three recurring sources: semantic ambiguity, lack of actionability, and context dependence. We further explore whether incorporating this taxonomy into LLM-based filtering improves the identification of problematic training instances, observing that detecting misaligned training pairs remains challenging. Based on these observations, we argue that improving review comment generation requires more than dataset cleaning alone, motivating explicit validity criteria, richer contextual inputs, and evaluation practices aligned with review intent and actionability.
Large Language Models (LLMs) often generate natural-language comments while writing code, and these comments become part of the context used to generate the code that follows. However, it remains unclear which properties of comments affect code-generation performance. We study this question through observational analys...
Da Pan, Zhensu Sun, Cenyuan Zhang et al.· 0 citations
A paired-prompt benchmark for human-versus-machine detection across English text, Python code, and mixed text–code documents shows that reliable deployment requires cross-domain evaluation, mixed-content testing, and calibration beyond in-distribution accuracy.
A high-quality benchmark of 1,000 code refinement instances from 328 Python, Java, and JavaScript repositories that focused on one of the most challenging code refinement scenarios that strictly requires repository-level knowledge reasoning, and a straightforward method, RepoRefiner, which retrieves repository-level co...
Ke Wang, Peng Lan, Jia-Kun Liu et al.· ACM Transactions on Software...· 1 citation
This work proposes SurveyReview, a reviewer-aligned, multi-dimensional benchmark and dataset for survey evaluation, and develops a strong baseline evaluator that substantially improves alignment with human reviewers, providing a competitive reference for future research.
Yuheng Zhang, Yuanchun Wang, Fanjin Zhang et al.· Proceedings of the 32nd ACM...· 0 citations
This work proposes Feedback-to-Rubrics, a problem setting for learning criteria from inline comments on artifacts, which infers rubrics from these comments and iteratively refines them by observing errors in comment prediction based on the inferred rubrics.
Kotaro Yoshida, So Kuroki, Yuki Imajuku et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.