Skip to content
Review

Intelligent automated code review and software quality evaluation: a survey of AutoML and large language model techniques

Aug 2026 · International Conference on Automated Software Engineering · Vol 33 · 0 citations · 63 references

TL;DR

This survey synthesises research from 2020 to 2026 on AI-driven techniques for code review and software quality evaluation, focusing on two complementary paradigms: Automated Machine Learning (AutoML) and Large Language Models (LLMs).

View source

Similar papers

Review Open access Jul 2026

A Systematic Literature Review on Automated Program Repair using Large Language Models

Current research is summarized to identify key gaps and future directions to optimize LLM based APR are proposed, to assure its reliability and scalability in real world software development.

Fatmaelzahra Hamdi, Ramadam Moawad, A. Mohsen · 0 citations
#software testing Preprint Aug 2026

An Empirical Evaluation of Using Large Language Models for Automated Model-Based Test Generation

This paper presents an empirical evaluation of Large Language Models (LLMs) for automated model-based test generation, compared with a state-of-the-art model-based testing tool (GraphWalker) and its built-in algorithms (random and quick random for edge and vertex coverage settings).

H. Şanli, Onur Kilinççeker, Cihat Çetinkaya · 0 citations
Review Sep 2026

Test-Driven Approaches to Software Engineering with Large Language Models: A Survey of Phases, Tasks, and Agent Skills

Tests increasingly participate in the decisions made by large language models and software engineering agents. They specify intended behavior, guide program construction and repair, select candidates, constrain transformations, and provide execution evidence for software analysis. These uses draw on test-driven development, yet differ substantially in test order, oracle availability, editable artifacts, and the role of execution. We present a structured scoping survey organized around the question of what decision a test changes. The review integrates 87 research and supporting records, with method- or protocol-level extraction for 83 records, alongside a separate collection of five practice resources. We distinguish the Red--Green--Refactor cycle from test-conditioned generation, execution-guided refinement, test-mediated analysis, and evaluation-only testing. We then compare code generation, repair, translation, refactoring, clone detection, code search, localization, training-data construction, and formal-specification validation. A dedicated analysis examines how agent workflows and reusable skills encode testing procedures and how their effects are evaluated. Across these tasks, the evidence supports treating test availability, test validity, feedback use, and evaluation independence as separate properties. Test passing alone does not establish behavioral equivalence, effective feedback, or process adherence; aggregate improvements can also conceal different outcomes across models, tasks, and denominators. We synthesize these distinctions into a mechanism taxonomy, a cross-task comparison, and a protocol-sensitive evidence analysis, and identify research directions in oracle validation, causal evaluation, long-horizon maintenance, and reusable test-driven agent capabilities

Yun-Hao Liang, Cheng-Guang Gan, Rui-Xuan Ying et al. · 0 citations
Open access 2020

Machine Learning for Code Smell Detection and Resolution

The study highlights the potential of intelligent, data-driven approaches to enhance software quality and support continuous code improvement by leveraging code metrics, syntactic and semantic features, and historical refactoring data.

Emily Johnson · 0 citations
Review Open access Aug 2026

Conceptual Model of a Hybrid Automated AI-Based Code Review System

Ensuring code quality has become one of the important challenges in modern software development processes. As software systems grow, code complexity, architectural complexity, security risks, and the likelihood of technical debt accumulation also increase. Traditional code review processes, which mainly rely on human expertise and manual analysis, do not always ensure fast and consistent evaluation. This paper presents the concept of a Hybrid Automated AI-Based Code Review System, which combines static analysis, rule-based validation, project context processing, and the capabilities of Large Language Models (LLMs). The proposed approach considers code review not merely as a process of generating comments, but as a context-aware multi-stage process that includes code change assessment, risk identification, problem localization, and generation of explainable recommendations. The research methodology is based on the analysis of existing AI-based code review approaches and the development of a hybrid architectural model. The proposed model utilizes multi-source context, including code changes, project structure, code dependencies, static analysis results, testing information, and Pull Request descriptions. The main contribution of the research lies in the development of a conceptual model of a hybrid AI-based code review system that integrates traditional software analysis methods with AI-based contextual reasoning. The proposed approach aims to improve the accuracy, explainability, and practical applicability of code review systems while supporting, rather than replacing, human reviewers.

T. Todua, Giorgi Tsitlidze · 0 citations
#software testing Preprint Sep 2026

Enhancing Automated Unit Test Generation for NLP Libraries Using Large Language Models

LLMSuite is proposed, a hybrid test generation framework that integrates self-refinement prompting with class-level LLM reasoning into the search-based testing process and complements manually written test suites by exercising domain-specific behaviors that are often left untested.

Amirhossein Deljouyi, Annibale Panichella, Andy Zaidman · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.