Skip to content
Preprint

Assessing Behavioral Validation in UI Component Test Suites Using Inferred Metamorphic Relations

Aug 2026 · 0 citations · 47 references
Computer Science

TL;DR

An MR-based framework that uses inferred metamorphic relations (MRs) as an empirical behavioral reference, rather than a complete specification, for assessing UI component test suites and provides a complementary relation-level perspective for assessing behavioral validation in modern UI component testing.

Abstract

UI component libraries are commonly assessed using execution-based metrics such as statement and branch coverage, yet these metrics provide limited insight into whether tests verify the behavioral relations implied by component APIs and documentation. This paper presents an MR-based framework that uses inferred metamorphic relations (MRs) as an empirical behavioral reference, rather than a complete specification, for assessing UI component test suites. Given a component's source, documentation, and tests, the framework infers component-specific MRs using a UI-specific taxonomy, aligns tests with the inferred relations through hybrid deterministic and semantic analysis, and computes relation-level MR coverage metrics. We manually validate both the inferred MR space and the test--MR alignment. Our evaluation shows that existing test suites exercise substantially more behavioral relations than they explicitly validate: MR Cover remains between 42.5% and 47.6% across three LLM configurations and consistently below MR Touch. Most uncovered relations are weak-oracle cases, where behaviors are exercised but lack explicit behavioral validation. MR coverage also complements execution-based coverage by revealing behavioral gaps not reflected by statement or branch coverage alone. We further assess practical relevance through issue-description mapping, oracle strengthening, and MR-relevant injected faults. Most reported issue descriptions can be mapped to inferred MR relation types; weak-oracle relations often expose missing validation evidence; and MR labels show a trend in MR-relevant fault detection. Overall, MR coverage provides a complementary relation-level perspective for assessing behavioral validation in modern UI component testing.

View source

Similar papers

Preprint Sep 2026

Evaluating the effectiveness of class-level LLM-generated test suites in Python

Reliable assessment of LLM-generated tests should treat executability as a gate and combine coverage with mutation testing and structural quality indicators, and in practice, model selection should precede prompt tuning.

Bilal Al-Ahmad, M. Harshvardhan, Khaled El-Fakih et al. · 0 citations
#software testing Book Open access Oct 2026

SPINACH: Inferring Properties of Web Applications for Property-Based Testing

The paper asks whether high-level conceptual specifications improve LLM-generated PBT quality and whether they help developers extend and maintain AI-generated systems (RQ2), and reports preliminary results applying Spinach to two open-source applications.

Savitha Ravi, Michael Coblenz · 0 citations
#artificial intelligence Preprint Sep 2026

Can LLMs Reason About Runtime Behavior? A Repository-Level Dynamic Benchmark

SWE-Flux is introduced, a repository-level benchmark for dynamic execution reasoning containing 480 execution-grounded instances across 12 real Python repositories, with gold answers automatically harvested from instrumented test executions rather than written manually or judged by LLMs.

Hamed Taherkhani, Mohammad Abdollahi, Melika Sepidband et al. · 0 citations
Open access Oct 2026

CAST: A Compiler-Based Framework for Systematically Testing LLM Compositional Safety

Large language models (LLMs) are increasingly used in software pipelines, raising concerns about harmful behaviors in security-critical domains. Existing safety evaluations predominantly probe models with single prompts or short interactions, and therefore do not capture how safety behaves under multi-step workflows wh...

Lu Yan, Zhuo Zhang, Xiang-Zhe Xu et al. · 0 citations
#software testing Preprint Sep 2026

How effective are traditional test criteria at detecting bugs in large language models generated code?

An empirical study involving 5 Large Language Models and 4 benchmarks evaluates the effectiveness and efficiency of 3 widely used adequacy criteria: statement coverage, branch coverage, and mutation testing, finding that mutation testing only marginally outperforms traditional coverage criteria in both triggering and d...

Asma Hamidi, Michael Konstantinou, R. Degiovanni et al. · 0 citations

Evaluation of large language models in api testing using open API documentation

Conventional automated REST API testing approaches often depend on rule-based logic, extensive configuration, or source-code access, which limits their adaptability in rapidly evolving development environments. Recent advances in Large Language Models (LLMs) offer new possibilities for automating API test generation th...

Phyu Thet Thet Kyaw · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.