An MR-based framework that uses inferred metamorphic relations (MRs) as an empirical behavioral reference, rather than a complete specification, for assessing UI component test suites and provides a complementary relation-level perspective for assessing behavioral validation in modern UI component testing.
Abstract
UI component libraries are commonly assessed using execution-based metrics such as statement and branch coverage, yet these metrics provide limited insight into whether tests verify the behavioral relations implied by component APIs and documentation. This paper presents an MR-based framework that uses inferred metamorphic relations (MRs) as an empirical behavioral reference, rather than a complete specification, for assessing UI component test suites. Given a component's source, documentation, and tests, the framework infers component-specific MRs using a UI-specific taxonomy, aligns tests with the inferred relations through hybrid deterministic and semantic analysis, and computes relation-level MR coverage metrics. We manually validate both the inferred MR space and the test--MR alignment. Our evaluation shows that existing test suites exercise substantially more behavioral relations than they explicitly validate: MR Cover remains between 42.5% and 47.6% across three LLM configurations and consistently below MR Touch. Most uncovered relations are weak-oracle cases, where behaviors are exercised but lack explicit behavioral validation. MR coverage also complements execution-based coverage by revealing behavioral gaps not reflected by statement or branch coverage alone. We further assess practical relevance through issue-description mapping, oracle strengthening, and MR-relevant injected faults. Most reported issue descriptions can be mapped to inferred MR relation types; weak-oracle relations often expose missing validation evidence; and MR labels show a trend in MR-relevant fault detection. Overall, MR coverage provides a complementary relation-level perspective for assessing behavioral validation in modern UI component testing.
Reliable assessment of LLM-generated tests should treat executability as a gate and combine coverage with mutation testing and structural quality indicators, and in practice, model selection should precede prompt tuning.
Bilal Al-Ahmad, M. Harshvardhan, Khaled El-Fakih et al.· 0 citations
The paper asks whether high-level conceptual specifications improve LLM-generated PBT quality and whether they help developers extend and maintain AI-generated systems (RQ2), and reports preliminary results applying Spinach to two open-source applications.
Savitha Ravi, Michael Coblenz· Proceedings of the 1st Inter...· 0 citations
SWE-Flux is introduced, a repository-level benchmark for dynamic execution reasoning containing 480 execution-grounded instances across 12 real Python repositories, with gold answers automatically harvested from instrumented test executions rather than written manually or judged by LLMs.
Hamed Taherkhani, Mohammad Abdollahi, Melika Sepidband et al.· 0 citations
Large language models (LLMs) are increasingly used in software pipelines, raising concerns about harmful behaviors in security-critical domains. Existing safety evaluations predominantly probe models with single prompts or short interactions, and therefore do not capture how safety behaves under multi-step workflows wh...
Lu Yan, Zhuo Zhang, Xiang-Zhe Xu et al.· Proceedings of the ACM on So...· 0 citations
An empirical study involving 5 Large Language Models and 4 benchmarks evaluates the effectiveness and efficiency of 3 widely used adequacy criteria: statement coverage, branch coverage, and mutation testing, finding that mutation testing only marginally outperforms traditional coverage criteria in both triggering and d...
Asma Hamidi, Michael Konstantinou, R. Degiovanni et al.· 0 citations
Conventional automated REST API testing approaches often depend on rule-based logic, extensive configuration, or source-code access, which limits their adaptability in rapidly evolving development environments. Recent advances in Large Language Models (LLMs) offer new possibilities for automating API test generation th...
Phyu Thet Thet Kyaw· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.