Skip to content
Review Open access

Trust-aware evaluation frameworks for large language model reliability in enterprise AI platforms

Jul 2026 · International Journal of Science and Research Archive · 0 citations

Abstract

Trust-aware evaluation is an emerging but rapidly consolidating research area for assessing the reliability of large language models in enterprise AI systems. Generative AI brings a new set of reliability challenges that extend beyond traditional accuracy metrics: An incorrect answer could affect key business decisions, customer interactions, knowledge management, security setups, or organizational accountability. This review summarizes peer-reviewed journal articles published between 2015 and 2025 related to key aspects of trust-aware LLM evaluation, including hallucination and factuality assessment, LLM evaluation methods, trustworthy AI governance, and human trust calibration. As indicated in the literature, current evaluation practice is still fragmented, both in terms of the various technical metrics and in terms of the documentation instruments, as well as on the interpretation of the data and the user-centred design of the trust cues, organizational governance. The most critical gaps are weak alignment between benchmark outcomes and enterprise risk, limited post-deployment monitoring, insufficient context-specific trust calibration, and limited validation of evaluation frameworks in operational platforms. The article argues for an evidence-based approach that connects model behaviour, platform controls, user reliance, and auditable governance in enterprise reliability assessment.

Read PDF