An RL-guided repair loop is used as an instrumented diagnostic environment to separate correct-candidate availability from candidate selection and establish general superiority of RL, a selector, or an output representation to single-file Python algorithmic repair on QuixBugs.
Abstract
Recent Large Language Models (LLMs) have driven an intuitive expectation that automated program repair (APR) can be treated as a direct code-rewriting task. However, practical APR systems require generated candidates to be machine-usable, safely validated, and appropriately ranked. This paper uses an RL-guided repair loop as an instrumented diagnostic environment to separate correct-candidate availability from candidate selection. We evaluate all 40 Python programs in QuixBugs using four code-LLM families, four versioned output-contract conditions, and five fixed random seeds. The resulting 3,200 candidate-generation conditions were evaluated with public tests for plausible correctness and a separately defined 1,200-case augmented oracle for semantic correctness. All 40 programs and all prespecified runs were retained. Correct candidates were available in 173 of 3,200 conditions and for 30 of 40 programs, showing that availability was a necessary upstream constraint under the evaluated conditions. Static greedy selection improved the program-level correct-outcome rate over first selection by 10.0 percentage points (95% CI [0.0, 23.3]; Holm-adjusted $p=.25$ ), while the separate sequential comparison improved by 5.0 points (95% CI [0.0, 12.5]; Holm-adjusted $p=.50$ ); neither effect was statistically established. These conclusions are limited to single-file Python algorithmic repair on QuixBugs and do not establish general superiority of RL, a selector, or an output representation.
This paper identifies patch verbosity as a major yet overlooked concern in LLM-based APR and proposes RECAP, a lightweight, plug-and-play adapter that attaches to existing repair frameworks after generation that achieves a substantially better size-correctness tradeoff.
Wen-Qiang Luo, J. Keung, Xiaoyu Shi et al.· 0 citations
Open-weight large language models (LLMs) can generate function-level programs from natural-language prompts, but plausible candidates still fail on hidden semantics and repeat mistakes across repair attempts. We present ReMCTS, an execution-grounded, memory-augmented, LLM-guided MCTS-style search framework. It organize...
Hui-Fei Wang, Xin-Yi Huang, Yi-Heng Sun et al.· 0 citations
Results support a focused conclusion: LLM-generated review is most useful as complementary semantic guidance when paired with deployment-oriented test selection, rather than as a standalone testing artifact.
Hui-Xiang Zhen, Zhihan Zhang· International Conference on...· 0 citations
Results suggest that causal-aware reasoning and stability-oriented design can improve the effectiveness of LLM-based APR, a causality-guided multi-agent repair framework that improves the repair stage of existing LLM-based localization pipelines.
Lei Yuan, Shaohua Liu, Yu Wang et al.· Empirical Software Engineeri...· 0 citations
This study empirically evaluates whether cost-efficient Large Language Models (LLMs) can be trusted to generate enterprise code to a written specification. Three models (Gemini Flash 3, GPT-5.4 mini and Claude Haiku 4.5) were asked to solve 992 algorithmic problems as Java Spring Boot service methods conforming to a ma...
The results show that point accuracy alone is insufficient for characterizing LLM reliability in assertion generation and motivate robustness-aware evaluation for AI-assisted hardware verification.
Fnu Aditi· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.