Jul 2026· Journal of Software and Systems Modeling· 1 citation· 22 references
TL;DR
The evaluation shows that MMT and PIT are significantly more effective and efficient than $$\mu $$ μ BERT and Jumble, and shows that further research in mutation operators, as enabled by MMT, has high potential.
Abstract
Mutation testing is an approach to checking the robustness of test suites. The program code is slightly modified by mutations to inject bugs, and a test suite is robust enough if it finds them. Mutation testing tools provide sets of mutation operators, such as swapping arithmetic operators, to make small modifications to the program. The results of mutation tests depend directly on the possible mutations. These mutations should cause actual changes in the program behavior, but also should not prevent the program from being loaded and executed. The more advanced mutations are, the more they challenge the test suite. Existing non-model-based mutation testing tools do not support the definition of advanced mutation operators that go beyond manipulating a small number of adjacent instructions within a single method. Thus, we present a
model-driven
approach where mutations of Java bytecode can be flexibly defined as model transformations. Our tool, Model-based Mutation Testing (MMT), implements this approach and includes model transformations for conventional and advanced mutation operators, such as deleting overridden methods or changing type casts. To evaluate the effectiveness and efficiency of model-driven mutation testing, we have applied MMT to all projects and versions in Defects4J, a well-established collection of real-world Java projects with reproducible bugs. We check for MMT’s ability to generate mutants close to real bugs and compare it with the non-model-based mutation testing tools Jumble, PIT, and
$$\mu $$
μ
BERT. Our evaluation shows that MMT and PIT are significantly more effective and efficient than
$$\mu $$
μ
BERT and Jumble. MMT even outperforms PIT in its ability to generate such realistic bugs, with similar efficiency per generated mutant. Jumble and
$$\mu $$
μ
BERT are one or two orders of magnitude slower than MMT and PIT. There are some bugs reconstructed by only one of the tools, including some that only the advanced operators of MMT could replicate. Fifteen percent of the Defects4J project versions had bugs that could not be reconstructed by the mutation operators of any of the investigated tools. This shows that further research in mutation operators, as enabled by MMT, has high potential. Our mutation testing tool MMT is available online https://gitlab.uni-marburg.de/fb12/plt/modbeam-mt/mmt, as well as all the evaluation data (Ancona et al., Evaluation data of comparison of mutation testing tools. https://doi.org/10.5281/zenodo.20054492).
An LLM-based, mutation testing-driven approach for test case generation by integrating the semantic understanding of large language models with the precise evaluation mechanism of mutation testing, paving a new path for intelligent test case enhancement.
Pei-Pei Yang, Kun Jia, Tian-Fang Ma et al.· International journal of sof...· 0 citations
An empirical study involving 5 Large Language Models and 4 benchmarks evaluates the effectiveness and efficiency of 3 widely used adequacy criteria: statement coverage, branch coverage, and mutation testing, finding that mutation testing only marginally outperforms traditional coverage criteria in both triggering and d...
Asma Hamidi, Michael Konstantinou, R. Degiovanni et al.· 0 citations
This paper extends the state-of-the-art FlowRepair approach by replacing a subset of its mutation operators with LLM-generated mutations, enabling more flexible and expressive patch generation and highlighting fundamental limitations of naively integrating LLMs into search-based APR.
Ayesha Irshad, Pablo Valle, J. Ayerdi et al.· 0 citations
It is found that the usefulness of coverage and mutation is highly context-dependent: in regression-style settings where the code provided to the LLM can be reasonably assumed bug-free, these metrics can provide meaningful signals when comparing across models; in another common scenario where the code-under-test may al...
The proposed metamorphic testing technique can effectively reveal faults that would remain undetected with pure fuzzing, and is introduced as a novel metamorphic testing technique tailored to transpilers.
Investigating the evolution of ModBEAM, a large-scale metamodel for Java bytecode, demonstrates how targeted metamodel evolution can improve both expressiveness and operational efficiency, and illustrates the benefits of automated quality assessment in guiding metamodel evolution.
F. Ancona, Philipp Wieber, Christoph Bockisch et al.· Proceedings of the ACM/IEEE...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.