Skip to content

BayesAME: Bayesian Active Model Evaluation

Jul 2026 · arXiv.org · Vol abs/2607.27023 · 0 citations · 37 references
Computer Science Mathematics

TL;DR

This comprehensive analysis addresses recent skepticism in the literature, establishing that non-random coreset selection is advantageous over random selection and highlighting that leveraging continuous response log-likelihoods over traditional binary scores significantly enhances estimation accuracy.

Abstract

Evaluating large generative models across benchmarks is time-consuming and computationally expensive. This drives the need for methods that can estimate full benchmark performance by evaluating models on only a subset of items, known as a coreset. Current literature mostly requires the practitioner to input a coreset size. However, when reliable performance estimation takes priority over efficiency, an evaluation method should also be capable of automatically determining a coreset size that reflects this priority. We introduce BayesAME, a sequential Bayesian framework specifically targeting automatic determination of the coreset size. BayesAME models performance as a random variable by defining a latent ability for each group of items sharing the same historical model performances, with a joint prior distribution encoding the belief that the target model behaves similarly to these historical models. The posterior distribution over these abilities is used to derive performance estimators, quantify performance uncertainty, and select items to add to the coreset via an information-gain criterion. The coreset is iteratively augmented until the performance estimate fluctuation and the performance uncertainty fall below their respective user-defined thresholds. We propose a multi-target extension that captures performance correlations across multiple target models to further reduce the coreset size. Through extensive experiments across diverse benchmarks, we demonstrate that BayesAME consistently outperforms sequential adaptations of existing methods. Crucially, our comprehensive analysis addresses recent skepticism in the literature, establishing that non-random coreset selection is advantageous over random selection. Finally, we highlight that leveraging continuous response log-likelihoods over traditional binary scores significantly enhances estimation accuracy.

View source

Similar papers

Preprint Aug 2026

Stochastic Bayes factors: why, when, and how

The Bayes factor (BF) is a central tool in Bayesian hypothesis testing and model selection, yet its practical use is often challenged. Classical BFs depend heavily on prior specification, cannot be applied with improper priors, and are typically interpreted through arbitrary evidence scales. Moreover, they fail to capt...

L. Egidi, I. Ntzoufras · 0 citations
#machine learning Preprint Aug 2026

Enhancing Bayesian Optimization and Active Learning Through Kernel Diversity

A unified framework, KENDO (Kernel ENsemble Disagreement-aware Operator), is proposed that integrates Ensemble Gaussian Processes (EGP) with disagreement-aware acquisition strategies and extends the approach to multi-objective optimization via random scalarization that preserves the single-optimizer conditioning struct...

Heng Zhang, Hao-Tian Xiang, Konstantinos D. Polyzos et al. · 1 citation
Preprint Aug 2026

Knowing When to Stop: Bayesian Optimal Stopping for LLM Evaluations

LLM evaluations often use fixed sampling budgets, testing every item the same number of times even after estimates are precise. We introduce optstop, a precision-based adaptive stopping framework that treats evaluation as a sequential measurement problem: keep sampling where uncertainty remains high, and stop where est...

Toby D. Pilditch · 1 citation
Review Aug 2026

Bayesian Inference Procedures for A/B Testing: An Overview

Bayesian inference for A/B testing is a family of prior and stopping-rule configurations with fundamentally different statistical properties, but it is often discussed as a single method, and no systematic overview exists. This paper organizes common configurations into a three-tier hierarchy: 1) posterior coherence wi...

M. Schultzberg, Mattias Frånberg · 0 citations
Review Sep 2026

From priors to performance: enhancing statistical efficiency with Bayesian dynamic borrowing

One of the main advantages of the Bayesian approach to statistical inference is the flexibility in incorporating information from various sources, from expert opinion to historical data. Whereas the literature on Bayesian dynamic borrowing is rich, practical guidance for how to implement such methods is comparatively s...

Xin-Xin Chen, E. Braga, Joseph G. Ibrahim et al. · 0 citations
Preprint Sep 2026

Empirical Bayes for compound adaptive experiments

We investigate Empirical Bayes (EB) methods in the context of compound adaptive experiments, where the arm distribution in each experiment follows a normal distribution with an unknown mean that we seek to estimate. There are two main EB strategies: $g$-modeling, which estimates the prior by maximizing the marginal lik...

Karun Adusumilli, Jia-Ying Gu, Jun-Fan Tao · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.