This paper extends the state-of-the-art FlowRepair approach by replacing a subset of its mutation operators with LLM-generated mutations, enabling more flexible and expressive patch generation and highlighting fundamental limitations of naively integrating LLMs into search-based APR.
Abstract
Search-based Automated Program Repair (APR) techniques rely on carefully designed mutation operators to explore the space of candidate fixes. Recent advances in Large Language Models (LLMs) suggest that generative models could replace such operators by dynamically proposing repairs. In this paper, we investigate this hypothesis in the context of Cyber-Physical Systems (CPSs) modeled in Simulink/Stateflow. We extend the state-of-the-art FlowRepair approach by replacing a subset of its mutation operators with LLM-generated mutations, enabling more flexible and expressive patch generation. We evaluate the approach on a benchmark of 19 real-world faulty Stateflow models across four CPS domains, using the same experimental setup as FlowRepair for controlled comparison under the same wall-clock budget. Contrary to expectations, in this controlled evaluation, the LLM-based mutation substantially degrades repair performance under the FlowRepair experimental setup. Across the tested LLM variants, the LLM-based repair produced plausible patches for 4-6 models and valid patches for 4 models, compared to 18 and 16, respectively, with the original approach. Our analysis reveals that, in this integration, LLMs struggle with precise symbolic edits, lack behavioral feedback, and generate a noisy search space that hinders effective exploration. Rather than showing a general limitation of LLMs for APR, these findings highlight fundamental limitations of naively integrating LLMs into search-based APR and motivate hybrid approaches that combine structured mutation with generative guidance.
Mutation testing evaluates test-suite adequacy by injecting synthetic faults into program code. However, traditional rule-based tools often generate large numbers of trivial, redundant, or equivalent mutants that limit their practical use for identifying gaps in a test suite. While recent large language model (LLM)-bas...
Nils Kiele, Zainab Saad, Zi-Rui Wang et al.· 0 citations
This paper identifies patch verbosity as a major yet overlooked concern in LLM-based APR and proposes RECAP, a lightweight, plug-and-play adapter that attaches to existing repair frameworks after generation that achieves a substantially better size-correctness tradeoff.
Wen-Qiang Luo, J. Keung, Xiaoyu Shi et al.· 0 citations
The results show that point accuracy alone is insufficient for characterizing LLM reliability in assertion generation and motivate robustness-aware evaluation for AI-assisted hardware verification.
An RL-guided repair loop is used as an instrumented diagnostic environment to separate correct-candidate availability from candidate selection and establish general superiority of RL, a selector, or an output representation to single-file Python algorithmic repair on QuixBugs.
This work presents SpecRead, a benchmark that isolates specification comprehension from generation ability, and is automatically scorable by deterministic checks, with gray-zone cases counted wrong under the conservative main scoring.
This paper proposes an approach that reformulates the generation problem as a code generation task, on which the LLMs excel, and evaluates the approach along four dimensions—scalability, consistency, diversity, and realism—across two use cases and three LLMs.
Javier Polo-Gambín, José A. Ruipérez-Valiente, José Antonio Hernández López· Proceedings of the ACM/IEEE...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.