Skip to content
Preprint

Hype Meets Reality: Large Language Models as Mutators in Search-based Automated Program Repair of Simulink-Stateflow Models

Aug 2026 · 0 citations · 34 references
Computer Science

TL;DR

This paper extends the state-of-the-art FlowRepair approach by replacing a subset of its mutation operators with LLM-generated mutations, enabling more flexible and expressive patch generation and highlighting fundamental limitations of naively integrating LLMs into search-based APR.

Abstract

Search-based Automated Program Repair (APR) techniques rely on carefully designed mutation operators to explore the space of candidate fixes. Recent advances in Large Language Models (LLMs) suggest that generative models could replace such operators by dynamically proposing repairs. In this paper, we investigate this hypothesis in the context of Cyber-Physical Systems (CPSs) modeled in Simulink/Stateflow. We extend the state-of-the-art FlowRepair approach by replacing a subset of its mutation operators with LLM-generated mutations, enabling more flexible and expressive patch generation. We evaluate the approach on a benchmark of 19 real-world faulty Stateflow models across four CPS domains, using the same experimental setup as FlowRepair for controlled comparison under the same wall-clock budget. Contrary to expectations, in this controlled evaluation, the LLM-based mutation substantially degrades repair performance under the FlowRepair experimental setup. Across the tested LLM variants, the LLM-based repair produced plausible patches for 4-6 models and valid patches for 4 models, compared to 18 and 16, respectively, with the original approach. Our analysis reveals that, in this integration, LLMs struggle with precise symbolic edits, lack behavioral feedback, and generate a noisy search space that hinders effective exploration. Rather than showing a general limitation of LLMs for APR, these findings highlight fundamental limitations of naively integrating LLMs into search-based APR and motivate hybrid approaches that combine structured mutation with generative guidance.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Beyond Rule-Based Mutation Testing: Test-Aware Mutant Generation Using Large Language Models

Mutation testing evaluates test-suite adequacy by injecting synthetic faults into program code. However, traditional rule-based tools often generate large numbers of trivial, redundant, or equivalent mutants that limit their practical use for identifying gaps in a test suite. While recent large language model (LLM)-bas...

Nils Kiele, Zainab Saad, Zi-Rui Wang et al. · 0 citations
Review Aug 2026

Refine After Generation: Toward Correct and Concise Patches in LLM-based Program Repair

This paper identifies patch verbosity as a major yet overlooked concern in LLM-based APR and proposes RECAP, a lightweight, plug-and-play adapter that attaches to existing repair frameworks after generation that achieves a substantially better size-correctness tradeoff.

Wen-Qiang Luo, J. Keung, Xiaoyu Shi et al. · 0 citations
Open access 2026

Diagnosing Candidate Quality Bottlenecks in RL-Guided LLM Code Repair

An RL-guided repair loop is used as an instrumented diagnostic environment to separate correct-candidate availability from candidate selection and establish general superiority of RL, a selector, or an output representation to single-file Python algorithmic repair on QuixBugs.

Jun-Jun Zhang, Giseop Noh · 0 citations
#software testing Book Open access Oct 2026

LLM-Based Instance Model Generation via Code Synthesis

This paper proposes an approach that reformulates the generation problem as a code generation task, on which the LLMs excel, and evaluates the approach along four dimensions—scalability, consistency, diversity, and realism—across two use cases and three LLMs.

Javier Polo-Gambín, José A. Ruipérez-Valiente, José Antonio Hernández López · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.