A Formal Trustworthiness Construct for Large Language Model-Based Test Generation: A Multidimensional Index Empirically Evaluated Through a Multi-Agent Study
Aug 2026· Electronics· Vol 15, pp. 3694· 0 citations· 28 references
TL;DR
This research formalises the trustworthiness of LLM-based unit test generation as a multidimensional index comprising reliability, hallucination resistance, maintainability, functional completeness, and human-reference alignment.
Abstract
Software code testing remains a critically important but labour-intensive process in software quality assurance. Existing research evaluates large language model (LLM)-based unit test generation using various quality metrics, such as correctness, coverage, mutation score, and test code smells. However, these single metrics do not reflect the trustworthiness of the unit test generation process. Therefore, this research formalises the trustworthiness of LLM-based unit test generation as a multidimensional index comprising reliability, hallucination resistance, maintainability, functional completeness, and human-reference alignment. In this research, we investigate the effect of prompt engineering strategies on the trustworthiness of LLM-generated unit tests and compare them with human-written tests for the same focal methods. Each dimension is fed by a distinct artefact-level measurement and grounded in dependability theory and ISO/IEC 25010:2023. A centralised multi-agent system generates, builds, repairs, and measures the tests, so that all inputs are collected automatically. The index is evaluated on real-world C# focal methods across 18 model × prompt configurations and a paired human-written baseline. The human baseline achieves the highest T-UTG value (0.904), and the best configuration, Combined × Gemini, achieves 0.788. Entropy weighting identifies maintainability and hallucination resistance as the most discriminating dimensions, and a rank-acceptability analysis over the whole weight simplex confirms that this ordering does not depend on the chosen weighting scheme.
This work presents the first cross-task empirical evaluation of LLMs spanning five RE-related activities, as well as replication materials supporting reproducibility, and a broader understanding of the capabilities, limitations, and practical readiness of current LLMs for RE.
Jacek Dabrowski, Manjeshwar Aniruddh Mallya, Alessio Ferrari et al.· 0 citations
These findings provide practical guidance for selecting LLM judges, designing role prompts, and employing multi-judge voting strategies in “automated software quality assurance”.
This paper investigates the reliability of LLMs in evaluating UML diagrams generated through reverse engineering processes (source code) and asks: do LLM assessments align with those of human experts?
Olena Chebanyuk, Carles Sierra· International Conference on...· 0 citations
A structured scoping survey organized around the question of what decision a test changes is presented, and distinguishes the Red--Green--Refactor cycle from test-conditioned generation, execution-guided refinement, test-mediated analysis, and evaluation-only testing.
Yun-Hao Liang, Cheng-Guang Gan, Rui-Xuan Ying et al.· 0 citations
A large-scale empirical study of quality assurance (QA) practices in 157 open-source LLM-based agent projects with at least 100 GitHub stars highlights the need to move beyond feature-level testing toward systematic end-to-end validation that ensures agent workflows remain within intended boundaries when interacting wi...
Wu-Yang Dai, Moses Openja, Jiho Shin et al.· 0 citations
LLM-based IDE agents are increasingly used to generate repository-grounded unit tests, yet common evaluations often rely on execution success or coverage. These metrics can miss deeper quality issues such as weak assertions, missing edge cases, poor isolation, and limited maintainability. This paper presents VibeCheck,...