The first empirical evaluation of Shaker for Python is presented, comparing non-order-dependent flaky tests from the ground-truth dataset of Gruber et al., and finding that Shaker provides no statistically significant detection advantage over plain re-execution.
Abstract
Flaky tests pass or fail non-deterministically on unchanged code, eroding trust in test suites and inflating the cost of every failure. Shaker detects them by injecting resource contention (CPU, memory, and I/O stress) to amplify non-determinism caused by concurrent execution, and was reported to detect 95% of the flaky tests in a Java and Android benchmark against 37.5% for plain re-execution (ReRun). We present the first empirical evaluation of Shaker for Python. Drawing non-order-dependent flaky tests from the ground-truth dataset of Gruber et al., we compare Shaker against a budget-matched ReRun baseline in a paired design, giving both techniques the same number of test executions: Each of 137 tests is run 100 times under each. As configured for Java and Android, Shaker provides no statistically significant detection advantage over plain re-execution (37.2% vs. 35.8%; McNemar exact p = 0.84). Two findings explain why. First, fewer than half of the ground-truth flaky tests reproduce as flaky at all on independent hardware under either technique, and most of the tests that fail to reproduce never diverge once across 100 runs. Second, the tests that do reproduce are dominated by flakiness from network interactions and randomness rather than the concurrency Shaker targets. Beyond the tool, this exposes a broader hazard for the field: reusing a flaky-test ground truth across execution environments silently converts genuine flaky tests into apparent true negatives, deflating any tool's measured recall.
Mutation analysis as an adequacy metric for kernel-benchmark oracles is introduced: deterministic rules inject 10 compilable faults into verified CUDA implementations of 188 KernelBench problems, 7{,}384 of them with an independent kill witness; any test protocol is scored by the fraction it detects.
Ming-Zhe Du, A. Lưu, Dong Huang et al.· 0 citations
Decompile-Diverge is proposed, a behavioral comparison oracle not relying on fixed or hand-crafted tests: for each function it synthesizes a driver, grows a fuzzing corpus from the reference, and reruns the decompiled code on the same inputs to detect changes in the function's behavior.
Chang Liu, Edward Raff, Kristopher K. Micinski· 1 citation
Large language models (LLMs) increasingly repair software vulnerabilities, but most evaluations judge only similarity to a developer fix or removal of the weakness, leaving no single judge trustworthy.
Patrick Deininger, Wolfgang Slany· Journal of Cybersecurity and...· 0 citations
Behavioral malware analysis relies on dynamic-analysis logs—sequences of Windows API calls recorded by sandboxes such as CAPEv2—assumed to represent the malware’s effects faithfully. To the best of our knowledge, this assumption has never been tested by re-executing the recorded calls. We propose WinAPIReplay, which re...
Property-based testing (PBT), introduced by Haskell’s Quickcheck, is becoming more popular with successful ports for other languages, such as Java’s junit-quickcheck. With PBT, developers write a property test and a data generator. The data generator takes a source of non-determinism and uses it to output well-formed d...
Jesse Coultas, Joseph Wiseman, Luís Pina· Proceedings of the ACM on So...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.