Skip to content

CollabEval: Statistically Efficient Collaborative Model Evaluation via Matrix Completion

Jul 2026 · arXiv.org · Vol abs/2607.05046 · 2 citations · 34 references
Computer Science

TL;DR

This work proposes Collaborative Evaluation (CollabEval), a simple, effective, and principled method for exploiting dependencies between historical runs of different models on the same tasks to improve statistical efficiency.

Abstract

Evaluating generative AI models is a routine, but resource-intensive, process that is conducted over and over again during the course of model development. In this work, we propose Collaborative Evaluation (CollabEval), a simple, effective, and principled method for exploiting dependencies between historical runs of different models on the same tasks to improve statistical efficiency. Specifically, our approach treats model evaluation as a matrix completion problem over an $M \times N$ matrix of evaluation scores, where $M$ is the total number of models and $N$ is the total number of evaluation prompts. We assume that a subset of these $M$ models are targeted for evaluation. For these target models only a small fraction, $p$, of prompts has been annotated with evaluation scores. Leveraging recent results in prediction-powered inference, we build a low-rank approximation of the score matrix, and use the reconstructed values as control variates in a manner that guarantees unbiased estimates of the true evaluation metric mean, in addition to statistically valid confidence intervals. Empirically, across a wide range of datasets, models, and sparsity levels $p$, we find that CollabEval substantially reduces the mean confidence interval size, and the mean squared error of the point estimate, compared to baseline methods at the same annotation budget.

View source

Similar papers

Jul 2026

BayesAME: Bayesian Active Model Evaluation

This comprehensive analysis addresses recent skepticism in the literature, establishing that non-random coreset selection is advantageous over random selection and highlighting that leveraging continuous response log-likelihoods over traditional binary scores significantly enhances estimation accuracy.

Paula Cordero Encinar, taylan. cemgil, Arnaud Doucet et al. · 0 citations
Preprint Aug 2026

The RAT: A Unified Bayesian Model for RAG Evaluation

Evaluating Retrieval-Augmented Generation (RAG) systems requires assessing not only end-to-end correctness but also how individual components interact and how errors propagate through the pipeline. We introduce a Bayesian evaluation framework that jointly models retrieval success, abstention behavior, and answer correc...

Pius von Däniken, Felix Matthias Saaro, Mark Cieliebak et al. · 0 citations
Preprint Aug 2026

MoPLEx: Estimating Plackett-Luce Mixture Models for Multi-Objective Alignment

The effectiveness of MoPLEx for tackling multi-way rankings following heterogeneous preferences through measuring alignment via gradients through measuring alignment via gradients is demonstrated.

Dongyue Li, Zi-Niu Zhang, Lu Wang et al. · 0 citations
Preprint Aug 2026

MISO: Model-Internal-State-Guided Optimization for Ranking Models

Model Internal State Optimization (MISO), a systems workflow that uses model internal states (MIS), including parameters, activations, gradients, and normalization statistics, to prioritize such local optimization decisions.

Yongzhen Zhang, Xiaoyu Deng, Yifan He et al. · 0 citations
#machine learning Preprint Aug 2026

Dynamically Allocating Evaluation Effort for Model Ranking

This work formalizes multi-model human evaluation as a best-arm identification problem in a multi-armed bandit setup with correlated arms, where pulling an arm corresponds to human-evaluating a model, and proves the optimality of the proposed algorithms and shows that it improves discrimination between top-performing m...

Vilém Zouhar, Julia Kreutzer, A. Lavie et al. · 0 citations
Jul 2026

Harnessing the Potential of Optimizing Data Mixtures via Bayesian Domain Reweighting

Experimental results demonstrate that proposed Bayesian domain weighting method could achieve stable and efficient domain weights learning, and identifies optimal mixtures while consuming substantially less data than search-based function-fitting methods, revitalizing optimization-based domain weighting for large-scale...

Xiang Yuan, Kai-Qing Lei, Zhenyu Jin et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.