It is concluded that instability should be treated as a standard evaluation axis in SE optimization, which should be routinely measured, reported alongside performance, and used to calibrate trust in any single run.
Abstract
In software analytics, rerunning the same analysis twice often yields different models and conclusions. This reduces trust in the model and limits its use. We find that model instability is a major problem. Across 127 multi-objective SE optimization problems (12,700 test cases), repeated runs of a state-of-the-art optimizer agree on only 13.7% of test cases, even under improved settings. We argue that this instability is not merely noise to tolerate, but a property that can be measured and managed. By adjusting how labels are spent, how complex the models become, and how splits are scored, we obtain models that agree 4.8 times as often as the default configuration. The standard deviation of optimization error falls by 22% on average (mean std 17.4 to 13.6), while recommendation quality improves rather than degrades. In terms of quality, the refined settings are statistically top-ranked on 119 of 127 datasets, compared to 74 for the defaults. We then test causal and data-locality interventions and find that they help only partially, suggesting a residual stability floor. Our evidence suggests there are fundamental limits to stability set by the data itself (noise, scarce labels, proxy objectives, and the many near-equivalent models a dataset admits). We conclude that instability should be treated as a standard evaluation axis in SE optimization, which should be routinely measured, reported alongside performance, and used to calibrate trust in any single run. The methods in this paper provide a baseline against which future efforts to reduce SBSE instability can be judged. To support open science, we offer the following reproduction package: https://tinyurl.com/Model-Instability
A graded, multi-family, failure-aware framework for stress-testing reasoning models that exposes structure that an aggregate score hides and exposes consistent weaknesses across all models.
This work studies 24 open-weight monitors spanning nine pretraining lineages and a 29x range of detection skill, and reports that across six attacker models the gain result holds in all six, the agreement and cancellation results in five of six.
Standard evaluation of large language models is challenged by varying the token generation budget, i.e., the maximum tokens a model may produce, across seven levels, evaluating four models on three reasoning benchmarks, and finding three findings that argue for budget-conditioned evaluation protocols.
Rodrigo Guedes de Souza, Alison R. Panisson· 1 citation
This work evaluates three dense and mixture-of-experts models on BBQ and BBQ-V under seven conditions spanning batching, quantization, benchmark reduction, and their combinations, and compares accuracy, bias severity and prevalence, reasoning quality, subgroup behavior, subset-membership stability, runtime, and measure...
A. Elkady, Aravind Narayanan, Rehana Riaz et al.· 0 citations
A modular benchmarking framework centered on quantitative XAI quality metrics: fidelity, stability, sparsity, computational cost, and faithfulness gap, plus an explicit method for operating the framework end-to-end is presented, plus an explicit method for operating the framework end-to-end.
Jonathan Herrera Vasquez, Miguel Herrero Uceda· Revista de investigación mul...· 0 citations
Trustworthy A/B Patterns is a community replication effort to evaluate selected patterns at high statistical power, finding that even at this scale, there was not enough power for key business metrics, such as revenue per user and purchase conversion rate, within practical time horizons.
Ron Kohavi, Jakub Linowski, Lukas Vermeer et al.· Proceedings of the 32nd ACM...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.