It is proved that common-mode error is not identifiable from internal judge scores alone and proposed a bounded two-anchor Bernstein certificate for finite-search error and regret.
Abstract
Language-model judges and reward models enable scalable supervision, but finite optimization can exploit evaluator errors rather than improve response quality. We characterize this failure through the covariance geometry of evaluator ensembles. For calibrated judges, the ensemble mean retains common-mode error along the all-ones direction, whereas cross-judge disagreement captures only orthogonal error. Consequently, disagreement can be high despite robust aggregation, or low while shared response-dependent errors persist. We prove that common-mode error is not identifiable from internal judge scores alone. Under a joint sub-Gaussian model, we bound best-of-K selection overstatement and target-quality regret, extending the guarantees to predictably adaptive search under conditional calibration. The resulting search terms scale as the square root of log K and are asymptotically tight for Gaussian projected errors. We further show that noisy quality proxies introduce artificial rank-one covariance without changing disagreement, and propose a bounded two-anchor Bernstein certificate for finite-search error and regret. Fixed-seed Gaussian stress tests over 120 (J, rho, K) configurations and real-model audits validate the theory while revealing the limits of disagreement-based diagnostics under increasing search pressure.
We formalize slate recommendation as a randomized score learner followed by deterministic selection. First, an appropriately scoped differential-privacy guarantee passes through selection and its audit trace by post-processing. End-to-end privacy holds only when selector inputs are public or independent, previous priva...
Sam Urmian, Qin-Yi Liu, Mohammad Khalil· Proceedings of the 20th ACM...· 0 citations
This work examines how reachable behavior, optimization budgets, and persistent adaptation shape exposure to proxy error, and develops a comparative framework for reward hacking across these three optimization substrates: weights, selection, and text.
Private Inference-Time Pessimism (PrivITP) is introduced, which combines $\chi^2$-regularized rejection sampling with a two-phase Gaussian mechanism, and achieves ex-post $(\epsilon,\delta)$-DP with a privacy cost independent of the number of responses, cleanly decouples the regularization parameter from the privacy pa...
I. Jain, Nandini Bhattad, Sayak Ray Chowdhury· 0 citations
This work empirically evaluates the primary performance-profile estimator on MATH-500, ScienceQA, MMLU, WinoGrande, HellaSwag, TruthfulQA, GSM8K, and ARC and proves calibration-free identification, unbiasedness for grouped profiles, local oracle optimality within centered linear augmentations, and first-order adaptivit...
Zhi Zhang, Lingfeng Lyu, Yue Kang et al.· 0 citations
In verifier-style RLVR, group-relative optimization often treats advantage scale as an implementation detail. This paper separates two low-variance cases: sub-resolution jitter that should not become a preference signal, and credible but small cardinal gaps that should be learned without distorting KL calibration. We p...
Best-of-$N$ reranking draws independent candidates from a reference policy and selects the response maximal under a fixed, sample-independent strict total order on outcomes. The selected law may differ substantially from the reference in Kullback–Leibler divergence. Prior work introduced a bounded statistic depending o...
Yu-Tong Zhang, Yao-Ran Yang· IEEE Signal Processing Lette...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.