Skip to content
Review

Rethinking Training Data for Generating Code Review Comments

Jul 2026 · arXiv.org · Vol abs/2607.25851 · 0 citations · 18 references
Computer Science

TL;DR

It is argued that improving review comment generation requires more than dataset cleaning alone, motivating explicit validity criteria, richer contextual inputs, and evaluation practices aligned with review intent and actionability.

Abstract

Generating code review comments has become a prominent research direction in automated code review, commonly formulated as a text generation task over diff-comment pairs. Despite advances in learning-based approaches, generated review comments are often generic, weakly grounded, or non-actionable. Recent studies have also shown that review comment datasets contain noisy or unsuitable training instances, motivating LLM-based dataset cleaning approaches. In this paper, we argue that problematic training instances are not homogeneous and that some limitations stem from deeper issues in the task formulation itself. Through an empirical inspection of a widely used review comment dataset, we identify misaligned training pairs: instances where the relationship between the code change and the review comment does not provide a reliable learning signal for generating actionable review feedback from localized inputs. We derive a taxonomy of misalignment capturing three recurring sources: semantic ambiguity, lack of actionability, and context dependence. We further explore whether incorporating this taxonomy into LLM-based filtering improves the identification of problematic training instances, observing that detecting misaligned training pairs remains challenging. Based on these observations, we argue that improving review comment generation requires more than dataset cleaning alone, motivating explicit validity criteria, richer contextual inputs, and evaluation practices aligned with review intent and actionability.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Talking to Itself While Coding: What Makes Comments Help Code Generation?

Large Language Models (LLMs) often generate natural-language comments while writing code, and these comments become part of the context used to generate the code that follows. However, it remains unclear which properties of comments affect code-generation performance. We study this question through observational analys...

Da Pan, Zhensu Sun, Cenyuan Zhang et al. · 0 citations
Open access Aug 2026

Detecting AI-Generated Text and Code: An Empirical Study of Cross-Generator and Cross-Domain Generalization

A paired-prompt benchmark for human-versus-machine detection across English text, Python code, and mixed text–code documents shows that reliable deployment requires cross-domain evaluation, mixed-content testing, and calibration beyond in-distribution accuracy.

Neethika Alluri, Pardha Saradhi Varma Gottumukkala, H. Indukuri · 0 citations
Review Aug 2026

Code Refinement with Repository Context: How Far are We?

A high-quality benchmark of 1,000 code refinement instances from 328 Python, Java, and JavaScript repositories that focused on one of the most challenging code refinement scenarios that strictly requires repository-level knowledge reasoning, and a straightforward method, RepoRefiner, which retrieves repository-level co...

Ke Wang, Peng Lan, Jia-Kun Liu et al. · 1 citation
Book Open access Aug 2026

SurveyReview: A Reviewer-Aligned Benchmark for Survey Evaluators

This work proposes SurveyReview, a reviewer-aligned, multi-dimensional benchmark and dataset for survey evaluation, and develops a strong baseline evaluator that substantially improves alignment with human reviewers, providing a competitive reference for future research.

Yuheng Zhang, Yuanchun Wang, Fanjin Zhang et al. · 0 citations
Review

Feedback-to-Rubrics: Can We Extract Expert Criteria from Inline Comments?

This work proposes Feedback-to-Rubrics, a problem setting for learning criteria from inline comments on artifacts, which infers rubrics from these comments and iteratively refines them by observing errors in comment prediction based on the inferred rubrics.

Kotaro Yoshida, So Kuroki, Yuki Imajuku et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.