This work proposes Collaborative Evaluation (CollabEval), a simple, effective, and principled method for exploiting dependencies between historical runs of different models on the same tasks to improve statistical efficiency.
Abstract
Evaluating generative AI models is a routine, but resource-intensive, process that is conducted over and over again during the course of model development. In this work, we propose Collaborative Evaluation (CollabEval), a simple, effective, and principled method for exploiting dependencies between historical runs of different models on the same tasks to improve statistical efficiency. Specifically, our approach treats model evaluation as a matrix completion problem over an $M \times N$ matrix of evaluation scores, where $M$ is the total number of models and $N$ is the total number of evaluation prompts. We assume that a subset of these $M$ models are targeted for evaluation. For these target models only a small fraction, $p$, of prompts has been annotated with evaluation scores. Leveraging recent results in prediction-powered inference, we build a low-rank approximation of the score matrix, and use the reconstructed values as control variates in a manner that guarantees unbiased estimates of the true evaluation metric mean, in addition to statistically valid confidence intervals. Empirically, across a wide range of datasets, models, and sparsity levels $p$, we find that CollabEval substantially reduces the mean confidence interval size, and the mean squared error of the point estimate, compared to baseline methods at the same annotation budget.
This comprehensive analysis addresses recent skepticism in the literature, establishing that non-random coreset selection is advantageous over random selection and highlighting that leveraging continuous response log-likelihoods over traditional binary scores significantly enhances estimation accuracy.
Paula Cordero Encinar, taylan. cemgil, Arnaud Doucet et al.· arXiv.org· 0 citations
Evaluating Retrieval-Augmented Generation (RAG) systems requires assessing not only end-to-end correctness but also how individual components interact and how errors propagate through the pipeline. We introduce a Bayesian evaluation framework that jointly models retrieval success, abstention behavior, and answer correc...
Pius von Däniken, Felix Matthias Saaro, Mark Cieliebak et al.· 0 citations
The effectiveness of MoPLEx for tackling multi-way rankings following heterogeneous preferences through measuring alignment via gradients through measuring alignment via gradients is demonstrated.
Dongyue Li, Zi-Niu Zhang, Lu Wang et al.· 0 citations
Model Internal State Optimization (MISO), a systems workflow that uses model internal states (MIS), including parameters, activations, gradients, and normalization statistics, to prioritize such local optimization decisions.
Yongzhen Zhang, Xiaoyu Deng, Yifan He et al.· 0 citations
This work formalizes multi-model human evaluation as a best-arm identification problem in a multi-armed bandit setup with correlated arms, where pulling an arm corresponds to human-evaluating a model, and proves the optimality of the proposed algorithms and shows that it improves discrimination between top-performing m...
Vilém Zouhar, Julia Kreutzer, A. Lavie et al.· 0 citations
Experimental results demonstrate that proposed Bayesian domain weighting method could achieve stable and efficient domain weights learning, and identifies optimal mixtures while consuming substantially less data than search-based function-fitting methods, revitalizing optimization-based domain weighting for large-scale...
Xiang Yuan, Kai-Qing Lei, Zhenyu Jin et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.