Skip to content
Open access

Characterizing LLM performance via a Bayesian lens

Jul 2026 · Frontiers in Applied Mathematics and Statistics · Vol 12 · 0 citations · 55 references

TL;DR

A hierarchical Bayesian Beta-Binomial framework for uncertainty-aware evaluation of LLMs on multiple-choice datasets, providing deeper insights into accuracy, consistency, and heterogeneity essential for reliable model benchmarking.

Abstract

While large language models (LLMs) are having a transformative impact on human society, evaluating them remains challenging. Standard benchmarks usually rely on single point estimates that obscure response stochasticity and variability in question difficulty. Here, we introduce a hierarchical Bayesian Beta-Binomial framework for uncertainty-aware evaluation of LLMs on multiple-choice datasets. Our approach moves beyond single accuracy metrics by modeling correct responses binomially and decomposing performance variation into intra-question stochasticity (response variability for a given question) and inter-question heterogeneity (variation in difficulty across questions) using separate priors. The framework yields a probabilistic assessment, providing full posterior distributions and credible intervals for mean accuracy, inter-question heterogeneity, and mean intra-question response variability, enabling rigorous uncertainty quantification. We demonstrate its utility by evaluating multiple LLMs across diverse benchmarks, including under semantic perturbations like question rephrasing. This analysis reveals nuanced model robustness insights and uncovers distinct behaviors across model classes (e.g., reasoning vs. non-reasoning) often missed by traditional descriptive statistics. This methodology offers a statistically grounded and powerful Bayesian lens for analyzing LLM performance, providing deeper insights into accuracy, consistency, and heterogeneity essential for reliable model benchmarking.

Read PDF

Similar papers

Preprint Aug 2026

LLM Evaluation on Unseen Questions: Contextual Multidimensional IRT Model

A model-based evaluation framework that combines multidimensional item response theory model with question contexts to predict LLM performance on unseen questions is studied, suggesting that context-aware psychometric modeling is a promising direction for efficient and interpretable LLM evaluation and highlighting cros...

Er-Gan Shang, Weijing Tang, Yinqiu He · 1 citation
Preprint Aug 2026

The RAT: A Unified Bayesian Model for RAG Evaluation

Evaluating Retrieval-Augmented Generation (RAG) systems requires assessing not only end-to-end correctness but also how individual components interact and how errors propagate through the pipeline. We introduce a Bayesian evaluation framework that jointly models retrieval success, abstention behavior, and answer correc...

Pius von Däniken, Felix Matthias Saaro, Mark Cieliebak et al. · 0 citations
Jul 2026

Laplace-PSN-IRT: Uncertainty Quantification for Neural Item Response Theory Models of LLM Benchmarks

Item Response Theory (IRT) has recently been proposed as a framework for evaluating large language model (LLM) benchmarks by separating a model's latent ability from the properties of individual benchmark items. Existing neural IRT approaches, including PSN-IRT, estimate these quantities using point estimates, limiting...

J. Francisco, Mandujano Reyes · 0 citations
Preprint Aug 2026

Asymptotic Risk Calibration for Selective Question Answering

A-CRC-QA is a post-hoc calibration framework for uncertainty-aware selective question answering that reformulates selection-conditioned error control as a linear expectation constraint and applies a monotonized empirical-risk calibration procedure inspired by conformal risk control.

Shufan Lin, Sijin Dong · 0 citations
#artificial intelligence Preprint Sep 2026

Pinocchio: Fast Uncertainty Estimates for Black-Box Language Models

In high-stakes decision-making applications of large language models (LLMs), practitioners require not only accurate LLMs but also uncertainty estimates for their predictions. Existing approaches to uncertainty estimation for LLMs require access to log-probabilities output by the model or require fine-tuning access. Ho...

K. D. Hayes, Arka Pal, Hao-Song Zhang et al. · 0 citations
Preprint Aug 2026

No Universal Signal Predicts Sample-Level LLM Regression under Version Updates

It is found that signal effectiveness is task-dependent: confidence is strongest on MCQ and simpler math, while likelihood/KL signals give the most frequent gains on harder math and code; no signal is universally best across model updates either, and some cross-version signals stay informative even when confidence fail...

Jiang-li Sheng, Yiwei Lu · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.