Skip to content
Open access

Diagnosing Candidate Quality Bottlenecks in RL-Guided LLM Code Repair

2026 · IEEE Access · Vol 14, pp. 131460-131475 · 0 citations · 26 references

TL;DR

An RL-guided repair loop is used as an instrumented diagnostic environment to separate correct-candidate availability from candidate selection and establish general superiority of RL, a selector, or an output representation to single-file Python algorithmic repair on QuixBugs.

Abstract

Recent Large Language Models (LLMs) have driven an intuitive expectation that automated program repair (APR) can be treated as a direct code-rewriting task. However, practical APR systems require generated candidates to be machine-usable, safely validated, and appropriately ranked. This paper uses an RL-guided repair loop as an instrumented diagnostic environment to separate correct-candidate availability from candidate selection. We evaluate all 40 Python programs in QuixBugs using four code-LLM families, four versioned output-contract conditions, and five fixed random seeds. The resulting 3,200 candidate-generation conditions were evaluated with public tests for plausible correctness and a separately defined 1,200-case augmented oracle for semantic correctness. All 40 programs and all prespecified runs were retained. Correct candidates were available in 173 of 3,200 conditions and for 30 of 40 programs, showing that availability was a necessary upstream constraint under the evaluated conditions. Static greedy selection improved the program-level correct-outcome rate over first selection by 10.0 percentage points (95% CI [0.0, 23.3]; Holm-adjusted $p=.25$ ), while the separate sequential comparison improved by 5.0 points (95% CI [0.0, 12.5]; Holm-adjusted $p=.50$ ); neither effect was statistically established. These conclusions are limited to single-file Python algorithmic repair on QuixBugs and do not establish general superiority of RL, a selector, or an output representation.

Read PDF

Similar papers

Review Aug 2026

Refine After Generation: Toward Correct and Concise Patches in LLM-based Program Repair

This paper identifies patch verbosity as a major yet overlooked concern in LLM-based APR and proposes RECAP, a lightweight, plug-and-play adapter that attaches to existing repair frameworks after generation that achieves a substantially better size-correctness tradeoff.

Wen-Qiang Luo, J. Keung, Xiaoyu Shi et al. · 0 citations
#natural language process... Preprint Sep 2026

ReMCTS: Reflection-Enhanced Monte Carlo Tree Search for Code Generation

Open-weight large language models (LLMs) can generate function-level programs from natural-language prompts, but plausible candidates still fail on hidden semantics and repeat mistakes across repair attempts. We present ReMCTS, an execution-grounded, memory-augmented, LLM-guided MCTS-style search framework. It organize...

Hui-Fei Wang, Xin-Yi Huang, Yi-Heng Sun et al. · 0 citations
Aug 2026

CGARF: a causality-guided framework for reliable automated program repair

Results suggest that causal-aware reasoning and stability-oriented design can improve the effectiveness of LLM-based APR, a causality-guided multi-agent repair framework that improves the repair stage of existing LLM-based localization pipelines.

Lei Yuan, Shaohua Liu, Yu Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

An Empirical Evaluation of Cost-Efficient Large Language Models on Algorithmic Programming Tasks

This study empirically evaluates whether cost-efficient Large Language Models (LLMs) can be trusted to generate enterprise code to a written specification. Three models (Gemini Flash 3, GPT-5.4 mini and Claude Haiku 4.5) were asked to solve 992 algorithmic problems as Java Spring Boot service methods conforming to a ma...

Chandimal Adikari, Nandika Herath · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.