Skip to content
Open access

A statistical framework for reliability evaluation of large language models

Sep 2026 · Advances in Engineering Innovation · Vol 17, pp. 96-106 · 0 citations

TL;DR

A statistical framework for quantitatively assessing LLM reliability under multiple evaluation criteria and shows that the proposed framework can identify reliability differences between different LLMs is reasonably robust to variations in indicator weights.

Abstract

The growing use of Large Language Models (LLMs) in real-world applications has increased the need for reliable evaluation methods. This study proposes a statistical framework for evaluating LLM reliability by jointly considering model accuracy, hallucination rate, and cross-domain performance stability. Based on the TruthfulQA dataset, 300 test questions were selected using a stratified sampling strategy, and responses were generated by DeepSeek-V4-Flash and GPT-4o-mini. GPT-4o was then employed as an automated evaluator to assess the 600 generated responses. Model reliability was evaluated using three dimensions: Accuracy, Hallucination Rate, and Cross-Domain Stability. In particular, the proposed CDS metric quantifies the consistency of model performance across different knowledge categories. A Weighted Geometric Mean was further employed to construct an overall reliability score. The experimental results show that the proposed framework can identify reliability differences between different LLMs. DeepSeek-V4-Flash achieved an overall reliability score of 0.770, compared with 0.678 for GPT-4o-mini. Sensitivity analysis further shows that the model ranking remains unchanged under different weight configurations, indicating that the proposed framework is reasonably robust to variations in indicator weights. This study provides a statistical framework for quantitatively assessing LLM reliability under multiple evaluation criteria.

Read PDF

Similar papers

#natural language process... Preprint Sep 2026

When Consistency Does Not Mean Reliability: Evaluating Local LLM Judges Against Human Ratings

The results show that high self-consistency does not necessarily indicate high agreement with human judgments when using local LLMs as automatic judges, and highlight the need to evaluate both consistency and human alignment when using local LLMs as automatic judges.

Aakash Kumar Tiwari · 1 citation
#artificial intelligence Preprint Sep 2026

Domain-Specific Hallucination Detection in Large Language Models

Large language models generate fluent text that can contain unfaithful claims -- a phenomenon known as hallucination. We present a multi-signal detection pipeline combining fine-tuned DeBERTa-v3 classification, Monte Carlo (MC) Dropout uncertainty quantification, and temperature-scaled calibration for response-level ha...

Varun Teja Chundru, Debasmita Biswas · 0 citations
Open access Aug 2026

Assessing reliability of BERT-based models on question answering tasks

This study evaluates the reliability of four BERT-based models by assessing response stability under two conditions: internal model variations induced via Monte Carlo Dropout (MCD) and input perturbations through paraphrasing.

Pooja Yadav, Priyanka Harjule, Basant Agarwal et al. · 0 citations
Review Open access Sep 2026

Information retrieval reliability in large language models: a study of source verification

Large language models (LLMs) have come into widespread use in recent years across domains ranging from education and journalism to academic research and everyday information seeking. Their ability to produce fluent, coherent-sounding answers in natural language creates the impression that these answers are accurate and...

Harun Nabiyev · 0 citations
Review Open access Nov 2025

Evaluating Large Language Models as Tools to Navigate Researchers in Rapidly Evolving Research Landscapes: A Case Study in Cancer Drug Response Prediction

Large Language Models (LLMs) have emerged as promising tools for assisting researchers in automating and accelerating the synthesis of literature reviews. However, their reliability is a significant concern due to issues like factual inaccuracies and hallucinations. The key question is whether LLMs can reliably provide...

Elena A. Mourelatou, Ioannis Katakis · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.