Jul 2026· Frontiers in Applied Mathematics and Statistics· Vol 12· 0 citations· 55 references
TL;DR
A hierarchical Bayesian Beta-Binomial framework for uncertainty-aware evaluation of LLMs on multiple-choice datasets, providing deeper insights into accuracy, consistency, and heterogeneity essential for reliable model benchmarking.
Abstract
While large language models (LLMs) are having a transformative impact on human society, evaluating them remains challenging. Standard benchmarks usually rely on single point estimates that obscure response stochasticity and variability in question difficulty. Here, we introduce a hierarchical Bayesian Beta-Binomial framework for uncertainty-aware evaluation of LLMs on multiple-choice datasets. Our approach moves beyond single accuracy metrics by modeling correct responses binomially and decomposing performance variation into intra-question stochasticity (response variability for a given question) and inter-question heterogeneity (variation in difficulty across questions) using separate priors. The framework yields a probabilistic assessment, providing full posterior distributions and credible intervals for mean accuracy, inter-question heterogeneity, and mean intra-question response variability, enabling rigorous uncertainty quantification. We demonstrate its utility by evaluating multiple LLMs across diverse benchmarks, including under semantic perturbations like question rephrasing. This analysis reveals nuanced model robustness insights and uncovers distinct behaviors across model classes (e.g., reasoning vs. non-reasoning) often missed by traditional descriptive statistics. This methodology offers a statistically grounded and powerful Bayesian lens for analyzing LLM performance, providing deeper insights into accuracy, consistency, and heterogeneity essential for reliable model benchmarking.
A model-based evaluation framework that combines multidimensional item response theory model with question contexts to predict LLM performance on unseen questions is studied, suggesting that context-aware psychometric modeling is a promising direction for efficient and interpretable LLM evaluation and highlighting cros...
Evaluating Retrieval-Augmented Generation (RAG) systems requires assessing not only end-to-end correctness but also how individual components interact and how errors propagate through the pipeline. We introduce a Bayesian evaluation framework that jointly models retrieval success, abstention behavior, and answer correc...
Pius von Däniken, Felix Matthias Saaro, Mark Cieliebak et al.· 0 citations
Item Response Theory (IRT) has recently been proposed as a framework for evaluating large language model (LLM) benchmarks by separating a model's latent ability from the properties of individual benchmark items. Existing neural IRT approaches, including PSN-IRT, estimate these quantities using point estimates, limiting...
J. Francisco, Mandujano Reyes· arXiv.org· 0 citations
A-CRC-QA is a post-hoc calibration framework for uncertainty-aware selective question answering that reformulates selection-conditioned error control as a linear expectation constraint and applies a monotonized empirical-risk calibration procedure inspired by conformal risk control.
In high-stakes decision-making applications of large language models (LLMs), practitioners require not only accurate LLMs but also uncertainty estimates for their predictions. Existing approaches to uncertainty estimation for LLMs require access to log-probabilities output by the model or require fine-tuning access. Ho...
K. D. Hayes, Arka Pal, Hao-Song Zhang et al.· 0 citations
It is found that signal effectiveness is task-dependent: confidence is strongest on MCQ and simpler math, while likelihood/KL signals give the most frequent gains on harder math and code; no signal is universally best across model updates either, and some cross-version signals stay informative even when confidence fail...
Jiang-li Sheng, Yiwei Lu· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.