Skip to content
Open access

Assessing model-driven mutation testing of Java bytecode

Jul 2026 · Journal of Software and Systems Modeling · 1 citation · 22 references

TL;DR

The evaluation shows that MMT and PIT are significantly more effective and efficient than $$\mu $$ μ BERT and Jumble, and shows that further research in mutation operators, as enabled by MMT, has high potential.

Abstract

Mutation testing is an approach to checking the robustness of test suites. The program code is slightly modified by mutations to inject bugs, and a test suite is robust enough if it finds them. Mutation testing tools provide sets of mutation operators, such as swapping arithmetic operators, to make small modifications to the program. The results of mutation tests depend directly on the possible mutations. These mutations should cause actual changes in the program behavior, but also should not prevent the program from being loaded and executed. The more advanced mutations are, the more they challenge the test suite. Existing non-model-based mutation testing tools do not support the definition of advanced mutation operators that go beyond manipulating a small number of adjacent instructions within a single method. Thus, we present a model-driven approach where mutations of Java bytecode can be flexibly defined as model transformations. Our tool, Model-based Mutation Testing (MMT), implements this approach and includes model transformations for conventional and advanced mutation operators, such as deleting overridden methods or changing type casts. To evaluate the effectiveness and efficiency of model-driven mutation testing, we have applied MMT to all projects and versions in Defects4J, a well-established collection of real-world Java projects with reproducible bugs. We check for MMT’s ability to generate mutants close to real bugs and compare it with the non-model-based mutation testing tools Jumble, PIT, and $$\mu $$ μ BERT. Our evaluation shows that MMT and PIT are significantly more effective and efficient than $$\mu $$ μ BERT and Jumble. MMT even outperforms PIT in its ability to generate such realistic bugs, with similar efficiency per generated mutant. Jumble and $$\mu $$ μ BERT are one or two orders of magnitude slower than MMT and PIT. There are some bugs reconstructed by only one of the tools, including some that only the advanced operators of MMT could replicate. Fifteen percent of the Defects4J project versions had bugs that could not be reconstructed by the mutation operators of any of the investigated tools. This shows that further research in mutation operators, as enabled by MMT, has high potential. Our mutation testing tool MMT is available online https://gitlab.uni-marburg.de/fb12/plt/modbeam-mt/mmt, as well as all the evaluation data (Ancona et al., Evaluation data of comparison of mutation testing tools. https://doi.org/10.5281/zenodo.20054492).

Read PDF

Similar papers

LLM-Mestcase: A Test Case Modification Method Based on LLMs and Mutation Testing

An LLM-based, mutation testing-driven approach for test case generation by integrating the semantic understanding of large language models with the precise evaluation mechanism of mutation testing, paving a new path for intelligent test case enhancement.

Pei-Pei Yang, Kun Jia, Tian-Fang Ma et al. · 0 citations
#software testing Preprint Sep 2026

How effective are traditional test criteria at detecting bugs in large language models generated code?

An empirical study involving 5 Large Language Models and 4 benchmarks evaluates the effectiveness and efficiency of 3 widely used adequacy criteria: statement coverage, branch coverage, and mutation testing, finding that mutation testing only marginally outperforms traditional coverage criteria in both triggering and d...

Asma Hamidi, Michael Konstantinou, R. Degiovanni et al. · 0 citations
Preprint Aug 2026

Hype Meets Reality: Large Language Models as Mutators in Search-based Automated Program Repair of Simulink-Stateflow Models

This paper extends the state-of-the-art FlowRepair approach by replacing a subset of its mutation operators with LLM-generated mutations, enabling more flexible and expressive patch generation and highlighting fundamental limitations of naively integrating LLMs into search-based APR.

Ayesha Irshad, Pablo Valle, J. Ayerdi et al. · 0 citations
Jul 2026

Do Coverage and Mutation Scores of LLM-Generated Test Suites Correlate with Their Effectiveness? (Replicability Study)

It is found that the usefulness of coverage and mutation is highly context-dependent: in regression-style settings where the code provided to the LLM can be reasonably assumed bug-free, these metrics can provide meaningful signals when comparing across models; in another common scenario where the code-under-test may al...

Junda Zhao, Shurui Zhou, Eldan Cohen · 0 citations
Jul 2026

Metamorphic Testing of Transpilers via Mutation Consistency of Programs

The proposed metamorphic testing technique can effectively reveal faults that would remain undetected with pure fuzzing, and is introduced as a novel metamorphic testing technique tailored to transpilers.

Enea Raffaele Ilario Papaleo, Luca Guglielmo, G. Denaro · 0 citations
Book Open access Oct 2026

Automated Quality Assessment of Metamodel Evolution: A Case Study in Java Mutation Testing

Investigating the evolution of ModBEAM, a large-scale metamodel for Java bytecode, demonstrates how targeted metamodel evolution can improve both expressiveness and operational efficiency, and illustrates the benefits of automated quality assessment in guiding metamodel evolution.

F. Ancona, Philipp Wieber, Christoph Bockisch et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.