Skip to content
Book Open access

Assessing Harmful Comments and Specificity in Code Review Feedback at Scale using Large Language Models

Jul 2026 · SIGSOFT FSE Companion · pp. 932-942 · 0 citations · 37 references
Computer Science

TL;DR

This study investigates how large language models can assess code review feedback quality along two dimensions, sentiment and specificity, to support more constructive collaboration, and demonstrates the feasibility and practical utility of automated feedback-quality assessment in real-world environments.

Abstract

Code review is central to collaborative software development, yet feedback quality can vary widely, influencing code maintainability and developer interactions. This study investigates how large language models (LLMs) can assess code review feedback quality along two dimensions, sentiment (with a focus on harmful comments) and specificity, to support more constructive collaboration. Using over 204,000 feedback threads from 30 open-source software (OSS) repositories, we evaluate eleven LLMs, achieving F1-scores up to 0.83 for sentiment and 0.67 for specificity. Most OSS feedback is neutral or low in specificity, with highly detailed or overtly harmful comments comprising a small minority. Industry data from 45 organisations contains significantly more highly specific feedback and more minimal reviews, while harmful feedback remains rare. Deployment of our approach in commercial settings demonstrated practical value. Specificity classifications delivered immediate value, such as revealing mentorship gaps when senior developers provided more specific feedback than they received, while harmful comment classifications required careful UX framing to avoid user sensitivity. Our findings demonstrate the feasibility and practical utility of automated feedback-quality assessment in real-world environments.

Read PDF

Similar papers

Review Aug 2026

Code Refinement with Repository Context: How Far are We?

A high-quality benchmark of 1,000 code refinement instances from 328 Python, Java, and JavaScript repositories that focused on one of the most challenging code refinement scenarios that strictly requires repository-level knowledge reasoning, and a straightforward method, RepoRefiner, which retrieves repository-level co...

Ke Wang, Peng Lan, Jia-Kun Liu et al. · 1 citation
Open access Jul 2026

Do influence tactics matter? investigating prompt framing effects in LLM code generation

Large Language Models (LLMs) are increasingly integrated into software engineering workflows, helping developers write, debug, test, and maintain code. While prompt wording and structure are known to influence model performance, the impact of psychologically inspired prompt framings remains unexplored. This study inves...

Alexandrina Deaconu, Anubhav Gupta, Manaal Basha et al. · 0 citations
Review Jul 2026

Rethinking Training Data for Generating Code Review Comments

It is argued that improving review comment generation requires more than dataset cleaning alone, motivating explicit validity criteria, richer contextual inputs, and evaluation practices aligned with review intent and actionability.

Leonardo Centellas-Claros, Estefania Pakarati-Cofre, Juan Pablo Sandoval Alcocer et al. · 0 citations
Preprint Aug 2026

Unreliable in Practice? A Comprehensive Study of Errors in LLM-Generated Code

It is observed that generated code often omits basic input validation or memory-safety checks, which can lead to overflows, resource exhaustion, or other reliability/security issues, and even the largest models frequently make simple mistakes.

Rodrigo Pato Nogueira, Marco Vieira, João R. Campos · 0 citations
#machine learning Preprint Sep 2026

On the Relation between Code Quality and Machine Learning Performance: A Large-scale Empirical Study

Context: Computational notebooks are the standard environment for machine learning (ML) development. Within the ML community, model performance is often the primary considered metric, and code quality is treated as a secondary concern. This prioritization relies on a largely untested assumption that code quality and ML...

Marius Mignard, Steven Costiou, Anne Etien · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.