Skip to content
Open access

A Formal Trustworthiness Construct for Large Language Model-Based Test Generation: A Multidimensional Index Empirically Evaluated Through a Multi-Agent Study

Aug 2026 · Electronics · Vol 15, pp. 3694 · 0 citations · 28 references

TL;DR

This research formalises the trustworthiness of LLM-based unit test generation as a multidimensional index comprising reliability, hallucination resistance, maintainability, functional completeness, and human-reference alignment.

Abstract

Software code testing remains a critically important but labour-intensive process in software quality assurance. Existing research evaluates large language model (LLM)-based unit test generation using various quality metrics, such as correctness, coverage, mutation score, and test code smells. However, these single metrics do not reflect the trustworthiness of the unit test generation process. Therefore, this research formalises the trustworthiness of LLM-based unit test generation as a multidimensional index comprising reliability, hallucination resistance, maintainability, functional completeness, and human-reference alignment. In this research, we investigate the effect of prompt engineering strategies on the trustworthiness of LLM-generated unit tests and compare them with human-written tests for the same focal methods. Each dimension is fed by a distinct artefact-level measurement and grounded in dependability theory and ISO/IEC 25010:2023. A centralised multi-agent system generates, builds, repairs, and measures the tests, so that all inputs are collected automatically. The index is evaluated on real-world C# focal methods across 18 model × prompt configurations and a paired human-written baseline. The human baseline achieves the highest T-UTG value (0.904), and the best configuration, Combined × Gemini, achieves 0.788. Entropy weighting identifies maintainability and hallucination resistance as the most discriminating dimensions, and a rank-acceptability analysis over the whole weight simplex confirms that this ordering does not depend on the chosen weighting scheme.

Read PDF

Similar papers

Preprint Aug 2026

Large Language Models for Requirements Engineering: A Cross-Task Empirical Evaluation

This work presents the first cross-task empirical evaluation of LLMs spanning five RE-related activities, as well as replication materials supporting reproducibility, and a broader understanding of the capabilities, limitations, and practical readiness of current LLMs for RE.

Jacek Dabrowski, Manjeshwar Aniruddh Mallya, Alessio Ferrari et al. · 0 citations
Open access 2026

A Comparative Study of LLMs and Human Judgment in UML Diagram Evaluation

This paper investigates the reliability of LLMs in evaluating UML diagrams generated through reverse engineering processes (source code) and asks: do LLM assessments align with those of human experts?

Olena Chebanyuk, Carles Sierra · 0 citations
Review Sep 2026

Test-Driven Approaches to Software Engineering with Large Language Models: A Survey of Phases, Tasks, and Agent Skills

A structured scoping survey organized around the question of what decision a test changes is presented, and distinguishes the Red--Green--Refactor cycle from test-conditioned generation, execution-guided refinement, test-mediated analysis, and evaluation-only testing.

Yun-Hao Liang, Cheng-Guang Gan, Rui-Xuan Ying et al. · 0 citations
#software testing Preprint Sep 2026

A Large-Scale Empirical Study of Quality Assurance Practices and Gaps in AI Agents

A large-scale empirical study of quality assurance (QA) practices in 157 open-source LLM-based agent projects with at least 100 GitHub stars highlights the need to move beyond feature-level testing toward systematic end-to-end validation that ensures agent workflows remain within intended boundaries when interacting wi...

Wu-Yang Dai, Moses Openja, Jiho Shin et al. · 0 citations
Preprint Sep 2026

VibeCheck: Assessing the Quality of LLM-Generated Unit Tests: A Multi-agent Empirical Study across Heterogeneous Repositories

LLM-based IDE agents are increasingly used to generate repository-grounded unit tests, yet common evaluations often rely on execution success or coverage. These metrics can miss deeper quality issues such as weak assertions, missing edge cases, poor isolation, and limited maintainability. This paper presents VibeCheck,...

Anika Tabassum, Mushahid Intesum, Md. Fahim Arefin et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.