Skip to content
Preprint

Evaluator Ensembles Under Reward Hacking: Covariance Geometry and Finite-Search Guarantees

Aug 2026 · 0 citations · 33 references
Computer Science

TL;DR

It is proved that common-mode error is not identifiable from internal judge scores alone and proposed a bounded two-anchor Bernstein certificate for finite-search error and regret.

Abstract

Language-model judges and reward models enable scalable supervision, but finite optimization can exploit evaluator errors rather than improve response quality. We characterize this failure through the covariance geometry of evaluator ensembles. For calibrated judges, the ensemble mean retains common-mode error along the all-ones direction, whereas cross-judge disagreement captures only orthogonal error. Consequently, disagreement can be high despite robust aggregation, or low while shared response-dependent errors persist. We prove that common-mode error is not identifiable from internal judge scores alone. Under a joint sub-Gaussian model, we bound best-of-K selection overstatement and target-quality regret, extending the guarantees to predictably adaptive search under conditional calibration. The resulting search terms scale as the square root of log K and are asymptotically tight for Gaussian projected errors. We further show that noisy quality proxies introduce artificial rank-one covariance without changing disagreement, and propose a bounded two-anchor Bernstein certificate for finite-search error and regret. Fixed-seed Gaussian stress tests over 120 (J, rho, K) configurations and real-model audits validate the theory while revealing the limits of disagreement-based diagnostics under increasing search pressure.

View source

Similar papers

#machine learning Book Open access Sep 2026

Decoupled Learning and Selection in Slate Recommendation for Privacy and Stability Under Noisy Scores

We formalize slate recommendation as a randomized score learner followed by deterministic selection. First, an appropriately scoped differential-privacy guarantee passes through selection and its audit trace by post-processing. End-to-end privacy holds only when selector inputs are public or independent, previous priva...

Sam Urmian, Qin-Yi Liu, Mohammad Khalil · 0 citations
Preprint Aug 2026

Privacy Without Regret: Differentially Private Inference-Time Alignment

Private Inference-Time Pessimism (PrivITP) is introduced, which combines $\chi^2$-regularized rejection sampling with a two-phase Gaussian mechanism, and achieves ex-post $(\epsilon,\delta)$-DP with a privacy cost independent of the number of responses, cleanly decouples the regularization parameter from the privacy pa...

I. Jain, Nandini Bhattad, Sayak Ray Chowdhury · 0 citations
Preprint Aug 2026

Conditional Evaluation of Language Models with Cheap Auxiliary Signals

This work empirically evaluates the primary performance-profile estimator on MATH-500, ScienceQA, MMLU, WinoGrande, HellaSwag, TruthfulQA, GSM8K, and ARC and proves calibration-free identification, unbiasedness for grouped profiles, local oracle optimality within centered linear augmentations, and first-order adaptivit...

Zhi Zhang, Lingfeng Lyu, Yue Kang et al. · 0 citations
#machine learning Preprint Aug 2026

Advantage Scale Calibration Imbalance in Group-Relative Optimization under Low-Variance Rewards: Diagnosis and Bounded Recovery

In verifier-style RLVR, group-relative optimization often treats advantage scale as an implementation detail. This paper separates two low-variance cases: sub-resolution jitter that should not become a preference signal, and credible but small cardinal gaps that should be learned without distorting KL calibration. We p...

Fei Ding · 0 citations
2026

A KL Certificate for Best-of-$N$ Reranking in Language-Model Inference

Best-of-$N$ reranking draws independent candidates from a reference policy and selects the response maximal under a fixed, sample-independent strict total order on outcomes. The selected law may differ substantially from the reference in Kullback–Leibler divergence. Prior work introduced a bounded statistic depending o...

Yu-Tong Zhang, Yao-Ran Yang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.